Course · 8 chapters

On-Call with AI

Fight ordinary production outages with AI agents: wire one into your telemetry, triage, debug live, verify the root cause, and decide exactly what it may touch in prod. All hands-on work runs on a provided lab.

Paidpractitioner8 chapters150 minEnglish + 6 languagesCertificate on completion

What you'll be able to do

  • Wire a read-only agent into your logs, metrics, and traces, then keep the same setup ready for Datadog or New Relic.
  • Turn a raw 3am page into a sourced, four-question incident brief in under two minutes.
  • Run a hypothesis loop with an agent against a real, scripted outage, and keep the playbook that finds it again.
  • Chase a bad deploy past its red herring to a cause a human can verify in five minutes.
  • Turn a messy prose runbook into steps an agent can run safely, with approval gates where it matters.
  • Write the policy that decides what your agent can read, what it can propose, and what it can never do alone.

What's inside

  1. 1
    On-Call with AI: Start Here

    Seven chapters, one lab, and a clear line between what your agent can read today and what it earns the right to touch later.

    10 min
  2. 2
    Your Stack, Agent-Ready

    Wire a read-only agent into your logs, metrics, and traces, then keep the same setup ready for Datadog or New Relic.

    20 min
  3. 3
    Alert Triage with an Agent

    Turn a raw 3am page into a sourced, four-question incident brief in under two minutes.

    20 min
  4. 4
    Live Debugging Under Pressure

    Run a hypothesis loop with an agent against a real, scripted outage, and keep the playbook that finds it again.

    20 min
  5. 5
    Root Cause Analysis That Holds Up

    Chase a bad deploy past its red herring to a cause a human can verify in five minutes.

    20 min
  6. 6
    Runbooks Agents Can Execute

    Turn a messy prose runbook into steps an agent can run safely, with approval gates where it matters.

    20 min
  7. 7
    Guardrails: What the Agent Touches in Prod

    Write the policy that decides what your agent can read, what it can propose, and what it can never do alone.

    20 min
  8. 8
    Postmortems & the Learning Loop

    Draft a blameless postmortem with an agent, file it with owned action items, and track whether the rotation is actually improving.

    20 min

Earn a certificate

Complete all chapters to receive your certificate of completion.