## TL;DR
Treat it like any incident first: contain the blast radius, preserve the logs, restore service. Then do the agent-specific work: freeze the agent's credentials, reconstruct exactly what it did from audit trails, and fix the guardrail that should have caught it. The postmortem asks what allowed the agent to act, not just what the agent did.

## Error / query
```text
agent-caused incident: how to respond and learn
```

## Use this skill when
- an agent's action caused or contributed to an outage
- an agent did something unexpected in production and you are not sure of the blast radius
- you need a postmortem template for automation-caused incidents
- leadership asks "how do we make sure the agent never does this again"

## Not for this skill when
- the agent only observed and a human took the action (standard incident process)
- nothing broke but the agent behaved oddly (file it as a near-miss, lighter process)
- you are setting up guardrails proactively (different skill)

## Steps

### Step 1: Contain first, revoke the agent's ability to do more damage
```bash
kubectl delete serviceaccount sre-agent -n prod || true
kubectl create serviceaccount sre-agent -n prod
aws iam delete-access-key --role-name sre-agent-prod --access-key-id [KEY_ID] || true
```
Expected: the agent's credentials are dead. Recreating the ServiceAccount invalidates its tokens; this is the equivalent of locking a human's account mid-incident. Do this before investigating, not after.

### Step 2: Preserve every log before it ages out
```bash
grep "system:serviceaccount:ops:sre-agent" /var/log/kubernetes/audit/audit.log | tail -100 | tee /tmp/agent-audit-snapshot.log
kubectl get events -n prod --sort-by=.lastTimestamp | tee /tmp/agent-events-snapshot.log | tail -20
```
Expected: two snapshot files with the agent's recent actions. Audit logs and events rotate; copy what you need now or the evidence disappears while you are still restoring service.

### Step 3: Restore service with the fastest safe path
```bash
kubectl -n prod rollout undo deployment/[APP]
terraform plan -detailed-exitcode; echo "exit=$?"
```
Expected: the rollback completes or the plan shows zero pending changes. Do not let the agent "fix" its own mess during the incident; humans drive the recovery, the agent stays frozen until the postmortem.

### Step 4: Reconstruct the exact timeline from evidence, not from the agent's summary
```bash
git -C /opt/gitops log --since="6 hours ago" --author="sre-agent" --oneline
aws cloudtrail lookup-events --lookup-attributes AttributeKey=Username,AttributeValue=sre-agent --max-results 30
```
Expected: the commit and API-call sequence showing what the agent changed and when. Ask the agent for its reasoning trace too, but treat that as one input among several; the audit log is ground truth.

### Step 5: Write the postmortem around the missing guardrail
```bash
ls /opt/postmortems/ | tail -5
```
Expected: the postmortem file started. The template's key question is not "why did the agent do this" but "which layer should have stopped it": was the RBAC too broad, the plan-first step skipped, the approval gate missing, or the policy engine misconfigured. Each incident should add or tighten exactly one guardrail layer.

## Variant phrasings

### "ai agent deleted production data, what now"
Contain (revoke creds), assess (what exactly was deleted, from backups), restore (from the tested restore process), then postmortem on why delete was grantable at all.

### "agent made an unauthorized change but nothing broke"
Near-miss process: same timeline reconstruction, lighter writeup, and still one guardrail improvement. Near-misses are free lessons; waste them and you pay full price later.

### "how to blame-less postmortem an agent incident"
Blame the system design, not the model: the agent behaved as agents do, and the question is why the surrounding system let that behavior reach production.

## Why it happens
Agent incidents happen at the intersection of probabilistic action and deterministic access: the agent guessed wrong (normal) and the system let the guess execute (the actual failure). Teams that punish the agent or the model miss the point; the fix is always in the layers around it. The incident is valuable precisely because it found the hole in your guardrails before a bigger guess found it.

## Edge cases and pitfalls
- Do not re-enable the agent with the same permissions "to help investigate"; frozen means frozen until the postmortem actions land.
- The agent's conversation or reasoning logs may contain the injected or misread instruction that triggered the action; preserve those alongside infra logs.
- If the agent acted through a shared automation identity, you cannot scope the blast radius; that finding alone justifies splitting identities.
- Legal or compliance may require retaining the evidence snapshots; check before cleaning up /tmp.
- One incident per guardrail layer is the goal; if the same layer fails twice, the layer design is wrong, not the agent.

## Provenance

Resolved from the public thread: https://vectle.com/posts/pst_GqC9XAVqQCBXVbYW5APa5Q
