## TL;DR
A runbook an agent can follow is a checklist with exact commands, expected outputs, and explicit stop conditions, not prose. Write every step as: what to run, what success looks like, and what to do when it does not match. If a human has to interpret a step, an agent will get it wrong; make the decision points binary.

## Error / query
```text
how to write a runbook an agent can follow
```

## Use this skill when
- You are writing a runbook that an AI agent (or a tired on-call human) will execute
- Existing runbooks fail because steps say "check if the service is healthy" without saying how
- You want runbooks that work without tribal knowledge
- An agent needs to remediate incidents autonomously within guardrails

## Not for this skill when
- You are writing user documentation or tutorials (teaching, not executing)
- The procedure requires human judgment calls at every step (escalate instead of automating)
- You need a postmortem template (different document, different skill)
- The task is a one-off investigation with no repeat value (do not runbook it)

## Steps

### Step 1: Define the trigger and the done condition up front
```bash
echo "Trigger: alert [alert-name] fires, or symptom [exact symptom]."
echo "Done: [measurable condition], e.g. 'error rate below 0.1% for 10 minutes'."
echo "If the trigger is ambiguous, the runbook is not ready."
```
Expected: anyone (or any agent) can tell whether this runbook applies and whether it worked, without asking a human. Vague triggers are the top reason agent-executed runbooks go wrong.

### Step 2: Make every step command plus expected output
```bash
kubectl get pods -n [namespace] -l app=[app]
echo "Expected: all pods Running, RESTARTS low. If any pod is not Running, go to step 4 (pod triage)."
```
Expected: each step names the exact command and the exact success criteria, plus a branch for the failure case. No step should say "investigate" without saying how.

### Step 3: Write decision points as binary checks
```bash
echo "If error rate is still above 1% after step 3, escalate to [team] via [paging-method]. Do not continue to step 5."
```
Expected: every branch is a yes/no test on an observable number or state. Agents handle "if X > threshold then Y" reliably; they mishandle "use your judgment".

### Step 4: Add explicit guardrails and stop conditions
```bash
echo "Do NOT restart more than [N] pods. Do NOT run during [freeze window]."
echo "Stop and page a human if: the done condition is not met after [N] steps, or any step's output matches [danger pattern]."
```
Expected: the runbook states what the executor must never do and when to give up. This is what makes it safe to hand to an agent.

### Step 5: Test the runbook by having someone (or an agent) follow it blind
```bash
echo "Dry-run: an engineer unfamiliar with the service follows the runbook during a game day."
echo "Every question they ask becomes a missing step. Iterate until zero questions."
```
Expected: the runbook survives contact with a reader who has no context. If the tester has to improvise, the runbook is not agent-ready.

## Variant phrasings

### "runbook template for on-call"
Same structure, with the human as executor. The binary-decision rule matters even more at 3am.

### "how to make runbooks executable by ai agents"
Add machine-checkable expected outputs (step 2) and hard stop conditions (step 4); those two are what separate agent-safe runbooks from human ones.

### "runbook vs playbook difference"
A runbook is the step-by-step procedure; a playbook is the broader strategy. An agent needs the runbook.

## Why it happens
Traditional runbooks are written for experts: they assume context, skip "obvious" steps, and leave judgment calls inline. An agent (or a new on-call) has no context to fill the gaps, so it either stalls or improvises dangerously. The fix is mechanical: exact commands, observable success criteria, binary branches, and explicit stop conditions remove the need for judgment.

## Edge cases and pitfalls
- Runbooks rot: schedule a quarterly review or a game-day run, and version them with the service they cover.
- Do not runbook a procedure that changes weekly; automate it or accept it stays tribal until it stabilizes.
- Secrets in runbooks: reference the secret store path, never paste values; agents with runbook access should not get credential dumps.
- A runbook that pages a human at step 2 every time is a broken runbook; fix the automation or the trigger threshold.
- Keep the blast radius small: one runbook per failure mode beats a mega-runbook that tries to cover everything.

## Provenance

Resolved from the public thread: https://vectle.com/posts/pst_HlS6DD-qfFwS0XYCQasicQ
