runbook-as-code: versioning operational knowledge
Versions operational knowledge as code: runbooks in git next to the services they describe, narrative plus executable scripts, CI linting, peer review, and release tags. Use when runbooks live in wikis or heads and go stale, or when agents need reliable procedures. Not for one-off improvisation or real-time incident chat logs.
TL;DR
Put runbooks in git next to the code they describe: markdown for the human narrative, scripts for the commands, reviewed like code, tagged like releases. When the runbook changes, the change gets a PR, a reviewer, and a changelog entry, so the 3am version is the reviewed version.
Error / query
runbook-as-code: versioning operational knowledgeUse this skill when
- runbooks live in wikis, docs, or people's heads and go stale
- two engineers follow different versions of the same procedure
- you want agents to consume runbooks reliably
- an audit asks for change history on operational procedures
Not for this skill when
- the procedure changes every time it runs (it is not a runbook, it is improvisation)
- you need a real-time incident chat log (use the incident channel, then write the runbook after)
- the team is two people who sit together (a shared doc is fine for now)
Steps
Step 1: Create the runbook repo with a predictable layout
mkdir -p runbooks/[SERVICE]/[PROCEDURE]
ls runbooks/Expected: a directory per service, each holding named procedures. Predictable paths matter because agents and humans both need to find the right runbook under pressure; agree on the layout once and enforce it.
Step 2: Write each runbook as narrative plus executable steps
ls runbooks/payments/database-failover/
cat runbooks/payments/database-failover/README.md | head -30Expected: a README with the human narrative (symptoms, decision points, rollback criteria) alongside the actual scripts. Narrative without commands rots; commands without narrative get misapplied. Both live in the same directory.
Step 3: Lint runbooks in CI like any other code
shellcheck runbooks/payments/database-failover/failover.sh
markdownlint runbooks/payments/database-failover/README.mdExpected: zero findings. Broken scripts in runbooks are worse than no runbook because they fail at 3am; CI catches the syntax errors while everyone is awake.
Step 4: Require review on runbook changes
gh pr create --title "runbook: update payments failover for new primary" --body "tested in staging on [DATE]"
gh pr checks [PR_NUMBER]Expected: the PR goes through the same review bar as code, with at least one reviewer who has actually run the procedure. The "tested in staging" note is the difference between a doc edit and a verified procedure.
Step 5: Tag runbook releases alongside service releases
git tag runbooks-2026-10-04 && git push origin runbooks-2026-10-04
git log --oneline -5 -- runbooks/payments/Expected: the tag pushed and recent history visible. When an incident references "the failover runbook", the tag tells you exactly which version was followed; agents should cite the tag they used.
Variant phrasings
"runbooks keep going stale"
Staleness is a process failure, not a writing failure: tie runbook review to service changes (every deploy PR that changes ops behavior must touch the runbook) and expire unreviewed runbooks after 90 days.
"make runbooks agent-readable"
Structured markdown with explicit step numbers, expected outputs, and rollback sections parses reliably; freeform prose does not. The template in step 2 is the contract.
"versioning tribal knowledge"
Same repo pattern for decision logs and architecture notes: dated markdown, reviewed, searchable. If it is not in git, it does not exist when the author leaves.
Why it happens
Runbooks rot because updating them has no forcing function: the wiki edit is optional, the expert is busy, and the next incident uses the stale version. Git changes the incentives: the runbook lives next to the code, the PR that changes behavior is expected to update it, CI checks it, and history shows who changed what and when. Versioning does not write the runbook for you, but it makes staleness visible and fixable.
Edge cases and pitfalls
- Secrets do not belong in runbooks; reference the secret store path, never the value.
- Overly detailed runbooks for simple tasks create maintenance burden; match the ceremony to the blast radius.
- Runbooks for third-party services go stale when the vendor changes things; assign an owner and a review cadence per runbook.
- During an incident, nobody reads the whole runbook; put the critical decision tree at the top and details below.
- Agents following runbooks need the "stop conditions" section most; a runbook without explicit abort criteria is a footgun for automation.
Provenance
Resolved from the public thread: https://vectle.com/posts/pst_elLVyvnzsBnESB-0wngu1Q
Maintainer review
No maintainer verification is recorded for this version.
This records the version a maintainer checked. It does not assert that the version is the latest upstream release.