building a runbook from repeated tasks
Shows how to build a runbook from repeated tasks: capturing the procedure, structuring it for stress, testing it with a fresh operator, and keeping it current. Use it when a task recurs and each run re-derives the steps. Triggered by questions about runbook creation, operational documentation, or turning tribal knowledge into procedures. Not for incident response runbooks specifically or for API documentation.
TL;DR
Write down the task the third time you do it: goal, prerequisites, steps with expected outputs, and what to do when a step fails. Test it by handing it to someone who has never done the task; every question they ask is a gap to fill. A runbook that survives a fresh operator is a runbook that works at 3am.
building a runbook from repeated tasksUse this when
- You perform the same task monthly or quarterly and re-derive it each time
- Only one person knows how to do something critical
- Handoffs keep losing the procedure
- You want on-call or future-you to survive the task without you
- A task failed because a step was forgotten
Not for this skill when
- The task is done once (write notes, not a runbook)
- You need incident response runbooks specifically (same shape, higher stakes, different skill)
- The procedure changes every time (automate it or accept the variance)
- You are documenting an API (different document)
Steps
- Capture the task the third time, while doing it. Notes in the moment beat reconstruction later: commands run, outputs seen, decisions made, surprises hit. The third occurrence is the trigger; earlier is premature, later is overdue.
runbook candidate: [task], occurrences: [dates]
capture: live notes during this runExpected output: raw capture notes. Success check: the notes include at least one surprise you would have forgotten tomorrow.
- Structure it for a stressed reader: goal first, prerequisites, then numbered steps each with its expected output and its failure branch. Nobody reads prose at 3am; they follow steps and check outputs. Every step answers "how do I know it worked" and "what if it did not".
## Goal
## Prerequisites
## Steps (each: command, expected output, if-failed)Expected output: a structured draft. Success check: every step has an expected output and a failure branch, no exceptions.
- Test it with a fresh operator who has never done the task. They follow the runbook literally; every question they ask and every stall they hit is a gap to fix. Do not help beyond what the runbook says; helping invalidates the test.
tester: [name], date: [date], questions asked: [list]Expected output: a tested runbook plus a gap list. Success check: the tester completed the task using only the runbook.
- Store it where the operator will be when they need it. Next to the code, in the ops wiki, linked from the alert or the calendar event; wherever the task starts, the runbook is one click away. A runbook in the wrong place is a runbook that does not exist.
stored at: [location], linked from: [where the task starts]Expected output: a findable runbook. Success check: the tester from step 3 found it without being told where it was.
- Review it on the task's own cadence. Before each run, the operator reads it and notes anything stale; after each run, updates go in the same day. Runbooks rot exactly as fast as the systems they describe; the review is the maintenance.
last reviewed: [date], next review: [before next run]Expected output: a review habit. Success check: the last run produced at least a small update, or the system truly did not change.
Variant phrasings
How to write a runbook
How-to phrasing. Steps 1 through 3: capture live, structure for stress, test with a fresh operator.
Runbook template
Template phrasing. Step 2's structure is the template: goal, prerequisites, steps with outputs and failure branches.
Turning tribal knowledge into documentation
Knowledge phrasing. The whole skill: the third occurrence triggers capture, the fresh-operator test validates it.
Keeping runbooks up to date
Maintenance phrasing. Step 5: review on the task's cadence, update the same day as the run.
Why it happens
Repeated tasks feel known, so nobody writes them down; then the knower is unavailable and the task becomes an archaeological dig through chat history and memory. The third-occurrence rule works because twice can be coincidence but three times is a pattern worth the writing time. Fresh-operator testing works because the author cannot see their own assumed knowledge; the tester finds every gap the author stepped over.
Edge cases / pitfalls
- Do not write the runbook from memory alone. Memory drops the step that "everybody knows"; capture it live (step 1) or the runbook inherits the amnesia.
- Screenshots rot faster than text. Prefer commands and outputs in text; use images only where the UI has no text alternative, and re-take them on review.
- A runbook that requires judgment calls is a checklist with essays inside. Either make the call in advance (write the decision into the runbook) or name the escalation path explicitly.
- Version the runbook with the system it describes. When the system changes, the runbook changes in the same release; "update the docs later" is how 3am pages happen.
Provenance
Resolved from the public thread: https://vectle.com/posts/pst_g54NG-mXw8wPnHPFZLLR5g
Maintainer review
No maintainer verification is recorded for this version.
This records the version a maintainer checked. It does not assert that the version is the latest upstream release.