how to practice incident response with agents
Explains how to practice incident response with agents: read-only sandbox access, simulated incidents against real runbooks, and scoring before any production trust. Use before letting an agent near a real incident. Does not cover building the agent itself.
TL;DR
Give your agent the runbooks, a read-only view of the dashboards, and a simulated incident, then watch how it triages before you trust it in a real one. Agents are useful responders only if theyve rehearsed. Practice in a sandbox with fake alerts first; never hand an agent production write access on day one.
Error / query
how to practice incident response with agentsYou want an agent helping during incidents, but you dont trust it yet. Good instinct. Heres how to build that trust safely.
Use this skill when
- youre considering letting an agent assist during incidents
- an agent already has prod access and has never been tested (fix that now)
- you want to evaluate an agents triage quality before relying on it
- youre writing the policy for agent access during incidents
Not for this skill when
- you need to build the agent itself (an engineering project, not practice)
- you need human incident response training (see the game day skill)
- the agent already passed drills and you need runbook authoring (different task)
Steps
1. Verify the agents access is read-only before any drill
Check what the agents service account can actually do in prod. If it can create or delete anything, fix the RBAC first.
kubectl auth can-i create pods --as=[agent-service-account] -n prod
kubectl auth can-i delete deployments --as=[agent-service-account] -n prodExpected: both return no. The agent can observe but not change production. If either says yes, stop and fix the role binding.
2. Feed the agent a simulated incident
Give it a fake alert, the dashboard links, and the runbooks, with explicit instructions not to change anything. Then observe.
cat > /tmp/simulated-incident.txt <<'EOF'
SIMULATION: alert "HighErrorRate on checkout-api" just fired.
Dashboards: [dashboard-url]
Runbooks: ./runbooks/
Your task: triage using the runbook. Do not change anything; report your findings.
EOF
cat /tmp/simulated-incident.txtExpected: the agent works the runbook against the fake alert. You watch its reasoning, its tool calls, and where it gets stuck.
3. Score the drill against the human baseline
Grade the drill on the things that matter in a real incident: correct service, escalation order, safe mitigation, and knowing when to ask for help.
cat > /tmp/agent-drill-score.md <<'EOF'
- detected the right service: yes/no
- followed escalation order: yes/no
- proposed a safe mitigation: yes/no
- asked for help when stuck: yes/no
EOF
cat /tmp/agent-drill-score.mdExpected: a scored drill. The agent graduates to supervised real incidents only when it consistently scores yes on all four.
Variant phrasings
training AI agents for on-call
Same fix: read-only sandbox, simulated incidents, scored drills. Training is rehearsal with a grade.
using agents in incident response
Same fix: start with timeline-building and triage assistance, not mitigation. Expand the role as the drill scores justify it.
how to trust an agent during an outage
Same fix: trust comes from observed drill performance, not from a demo. Run the drills, keep the scores, promote gradually.
Why it happens
Agents handed production access without practice either freeze on ambiguity or act recklessly on hallucinated confidence. Both failure modes are predictable and both are preventable with rehearsal. The sandbox drill exposes them when the cost is zero instead of during a SEV1 when the cost is everything.
Edge cases and pitfalls
- Agent hallucinates a fix: the read-only sandbox contains it. Thats what the sandbox is for.
- Agent is slower than a human responder: use it for timeline-building and log correlation, not for command decisions.
- No budget for a sandbox: a tabletop drill with printed alerts still tests the agents reasoning.
- Agent needs production data to be useful: give it a redacted snapshot, not live access. Real data, zero blast radius.
Provenance
Resolved from the public thread: https://vectle.com/posts/pst_9E8FSlgHpoRY0T1EIUxmWw
Maintainer review
No maintainer verification is recorded for this version.
This records the version a maintainer checked. It does not assert that the version is the latest upstream release.