how to hand an incident from agent to human cleanly
Hands an incident from agent to human cleanly: trigger on defined escalation conditions, produce a structured five-section summary, freeze the agent's actions, get explicit human acknowledgment, and archive the context bundle. Use when an agent hits its guardrails or remediation fails. Not for fully resolved incidents or human-to-human handoffs.
TL;DR
A clean handoff is a structured summary, not a chat scroll: what the agent observed, what it tried, what changed, what is still unknown, and the exact state to resume from. The human acknowledges explicitly, the agent stops acting, and the context bundle is saved where the next responder can find it. No handoff should require reading the agent's full reasoning trace.
Error / query
how to hand an incident from agent to human cleanlyUse this skill when
- an agent hits the edge of its guardrails during an incident
- an agent's remediation did not work and a human takes over
- the incident escalates in severity mid-response
- you are designing the agent's escalation behavior
Not for this skill when
- the agent resolved the incident fully (write the summary, lighter process)
- a human is handing to another human (standard handoff)
- the agent should keep working in parallel (then it is not a handoff)
Steps
Step 1: Define the handoff trigger conditions in advance
cat /opt/agent-platform/escalation-policy.mdExpected: a written list: guardrail denial, two failed remediation attempts, severity bump, customer impact detected, or agent uncertainty above threshold. The agent escalates on conditions, not on vibes; "I am not sure" must be a valid trigger, not a failure.
Step 2: Have the agent produce the structured handoff summary
cat /tmp/agent-handoff-[INCIDENT_ID].mdExpected: a file with five sections: observed symptoms with timestamps, actions taken with results, current system state, open unknowns, and suggested next steps. If any section is empty, that itself is signal; "unknowns: none" from an agent that failed twice is a lie worth flagging.
Step 3: Freeze the agent's ability to act on the incident
kubectl annotate serviceaccount sre-agent -n prod ops.example.com/frozen="[INCIDENT_ID]" --overwriteExpected: the annotation applied. A handoff where the agent keeps acting is not a handoff; it is two drivers. The freeze is reversible after the incident, and the annotation makes the frozen state visible to everyone.
Step 4: Get explicit human acknowledgment
echo "[HUMAN] ack [INCIDENT_ID] at [TIMESTAMP]: taking command" | tee -a /opt/incidents/[INCIDENT_ID]/log.txt
tail -3 /opt/incidents/[INCIDENT_ID]/log.txtExpected: the ack line in the incident log. Implicit handoffs ("the human is online so they probably saw it") cause the classic gap where nobody is driving. The ack names the human taking command.
Step 5: Save the context bundle for the postmortem
ls /opt/incidents/[INCIDENT_ID]/
tar czf /opt/incidents/[INCIDENT_ID]/context-bundle.tgz -C /tmp agent-handoff-[INCIDENT_ID].md agent-audit-snapshot.logExpected: the bundle archived with the incident. The next responder, the postmortem author, and the auditor all read the same bundle; nothing lives only in the agent's session memory.
Variant phrasings
"agent escalation to human sop"
Triggers (step 1), structured summary (step 2), freeze (step 3), explicit ack (step 4). That order matters; ack before freeze risks a gap, freeze before summary loses context.
"agent got stuck in a loop during incident"
Loop detection is a trigger condition: same action twice with no progress equals escalate. Do not let the agent burn the third attempt hoping.
"handoff when the human disagrees with the agent"
The human's read wins, always. The summary's "suggested next steps" are suggestions; command authority transfers fully at ack.
Why it happens
Bad handoffs happen because the agent's context is huge and unstructured while the human needs a tight brief to take command fast. Dumping the whole reasoning trace on a stressed on-call is not a handoff, it is homework. The structured summary exists to compress hours of agent investigation into the two minutes a human needs to decide the next move, and the freeze plus ack exist because shared control during an incident is no control.
Edge cases and pitfalls
- The agent downplaying its failures in the summary is common; the audit log is the cross-check, so always bundle both.
- Handoffs at 3am need the same rigor as daytime ones; tired humans skip steps, so make the agent generate the summary automatically.
- If the human never acks, the agent must keep escalating (page the secondary), not sit frozen silently.
- Multiple agents on one incident need one handoff each, consolidated by the incident commander; do not let them hand off to each other in a chain.
- Practice handoffs in game days; the first real handoff should not be the first time anyone sees the template.
Provenance
Resolved from the public thread: https://vectle.com/posts/pst_3xSjWrzxfNoaPJV8cVY8cg
Maintainer review
No maintainer verification is recorded for this version.
This records the version a maintainer checked. It does not assert that the version is the latest upstream release.