## TL;DR
A runbook that lives in a wiki rots; a runbook that lives next to the code gets updated in the same PR. Keep runbooks in the service repo (a `runbooks/` or `docs/runbooks/` directory), link them from the service README and from alerts, and require runbook updates in the PR checklist for operational changes. Versioning the runbook with the code means the rollback also rolls back the instructions.

## Error / query
```text
how to version runbooks alongside service code
```

## Use this skill when
- Runbooks are stale the moment an incident hits
- You want runbook changes reviewed with code changes
- Rollbacks should restore the matching operational docs
- Linking alerts and dashboards to the right runbook version

## Not for this skill when
- Writing the runbook content itself (authoring guide)
- Running the incident (incident response)
- Organizing a company wiki (knowledge management)

## Steps

### Step 1: Put runbooks in the service repo
```text
service-repo/
  runbooks/
    deploy.md
    database-failover.md
    oncall-quickstart.md
  README.md  -> links to runbooks/
```
Expected: the runbook files live next to the code they describe. Any checkout at a given commit has the matching runbook version; `git log` on the runbook shows when it last changed relative to the code.

### Step 2: Link runbooks from alerts and dashboards
```yaml
# in the alert annotation:
annotations:
  runbook_url: "https://git.YOUR-host/[org]/[repo]/blob/[branch]/runbooks/database-failover.md"
```
Expected: the alert links to the runbook at a branch or tag, not a wiki page. Pin the link to the deployed version's tag so the on-call sees the instructions matching production, not main.

### Step 3: Require runbook updates in code review
```text
PR template checklist:
- [ ] Operational change? Runbook updated in runbooks/
- [ ] New alert? runbook_url annotation added
- [ ] Removed feature? Its runbook section removed
```
Expected: reviewers catch stale runbooks the same way they catch missing tests. The checklist makes runbook maintenance part of the change, not a follow-up that never happens.

### Step 4: Verify runbooks stay fresh with a staleness check
```bash
git log --since="90 days ago" --name-only --pretty=format: -- runbooks/ | sort -u
```
Expected: the list of runbooks touched recently. Runbooks untouched in 6+ months while their service changed are stale by definition; schedule a review pass on the untouched ones.

## Variant phrasings

### "runbook as code"
Steps 1-3. Markdown in the repo, reviewed like code, linked from alerts.

### "keep runbooks up to date"
Co-location (step 1) plus the PR checklist (step 3) is the mechanism that actually works.

## Why it happens
Runbooks rot because updating them is a separate workflow from changing the system: different tool, different review, no trigger. Co-locating them with the code ties the update to the change that made it necessary, puts it in front of the same reviewers, and versions it with the deployment it describes.

## Edge cases and pitfalls
- Secrets in runbooks: keep credentials out; link to the secret store instead. Runbooks in git get cloned widely.
- Monorepos need per-service runbook dirs with clear ownership; a single top-level runbooks folder becomes nobody's job.
- Alert links to `main` drift from what is deployed; link to release tags for production alerts.
- Deleting a service should delete its runbooks; orphaned runbooks mislead the next incident.

## Provenance

Resolved from the public thread: https://vectle.com/posts/pst_awtnqG_qCsrRpt8sEIgGCQ
