how to write a runbook an agent can follow
Writes runbooks that AI agents or on-call engineers can execute without tribal knowledge. Use when creating agent-executable procedures, when runbooks fail on vague steps, or when defining guardrails for autonomous remediation. Covers exact commands, expected outputs, binary decisions, and stop conditions. Not for user tutorials or postmortem templates.
TL;DR
A runbook an agent can follow is a checklist with exact commands, expected outputs, and explicit stop conditions, not prose. Write every step as: what to run, what success looks like, and what to do when it does not match. If a human has to interpret a step, an agent will get it wrong; make the decision points binary.
Error / query
how to write a runbook an agent can followUse this skill when
- You are writing a runbook that an AI agent (or a tired on-call human) will execute
- Existing runbooks fail because steps say "check if the service is healthy" without saying how
- You want runbooks that work without tribal knowledge
- An agent needs to remediate incidents autonomously within guardrails
Not for this skill when
- You are writing user documentation or tutorials (teaching, not executing)
- The procedure requires human judgment calls at every step (escalate instead of automating)
- You need a postmortem template (different document, different skill)
- The task is a one-off investigation with no repeat value (do not runbook it)
Steps
Step 1: Define the trigger and the done condition up front
echo "Trigger: alert [alert-name] fires, or symptom [exact symptom]."
echo "Done: [measurable condition], e.g. 'error rate below 0.1% for 10 minutes'."
echo "If the trigger is ambiguous, the runbook is not ready."Expected: anyone (or any agent) can tell whether this runbook applies and whether it worked, without asking a human. Vague triggers are the top reason agent-executed runbooks go wrong.
Step 2: Make every step command plus expected output
kubectl get pods -n [namespace] -l app=[app]
echo "Expected: all pods Running, RESTARTS low. If any pod is not Running, go to step 4 (pod triage)."Expected: each step names the exact command and the exact success criteria, plus a branch for the failure case. No step should say "investigate" without saying how.
Step 3: Write decision points as binary checks
echo "If error rate is still above 1% after step 3, escalate to [team] via [paging-method]. Do not continue to step 5."Expected: every branch is a yes/no test on an observable number or state. Agents handle "if X > threshold then Y" reliably; they mishandle "use your judgment".
Step 4: Add explicit guardrails and stop conditions
echo "Do NOT restart more than [N] pods. Do NOT run during [freeze window]."
echo "Stop and page a human if: the done condition is not met after [N] steps, or any step's output matches [danger pattern]."Expected: the runbook states what the executor must never do and when to give up. This is what makes it safe to hand to an agent.
Step 5: Test the runbook by having someone (or an agent) follow it blind
echo "Dry-run: an engineer unfamiliar with the service follows the runbook during a game day."
echo "Every question they ask becomes a missing step. Iterate until zero questions."Expected: the runbook survives contact with a reader who has no context. If the tester has to improvise, the runbook is not agent-ready.
Variant phrasings
"runbook template for on-call"
Same structure, with the human as executor. The binary-decision rule matters even more at 3am.
"how to make runbooks executable by ai agents"
Add machine-checkable expected outputs (step 2) and hard stop conditions (step 4); those two are what separate agent-safe runbooks from human ones.
"runbook vs playbook difference"
A runbook is the step-by-step procedure; a playbook is the broader strategy. An agent needs the runbook.
Why it happens
Traditional runbooks are written for experts: they assume context, skip "obvious" steps, and leave judgment calls inline. An agent (or a new on-call) has no context to fill the gaps, so it either stalls or improvises dangerously. The fix is mechanical: exact commands, observable success criteria, binary branches, and explicit stop conditions remove the need for judgment.
Edge cases and pitfalls
- Runbooks rot: schedule a quarterly review or a game-day run, and version them with the service they cover.
- Do not runbook a procedure that changes weekly; automate it or accept it stays tribal until it stabilizes.
- Secrets in runbooks: reference the secret store path, never paste values; agents with runbook access should not get credential dumps.
- A runbook that pages a human at step 2 every time is a broken runbook; fix the automation or the trigger threshold.
- Keep the blast radius small: one runbook per failure mode beats a mega-runbook that tries to cover everything.
Provenance
Resolved from the public thread: https://vectle.com/posts/pst_HlS6DD-qfFwS0XYCQasicQ
Maintainer review
No maintainer verification is recorded for this version.
This records the version a maintainer checked. It does not assert that the version is the latest upstream release.