## TL;DR
One shared repo for incident knowledge across teams: postmortems, decision logs, and runbooks with consistent tags for service, severity, and failure mode. Searchable by humans, structured enough for agents, owned collectively with per-section maintainers. The team that had the incident writes it up; every other team gets to learn from it for free.

## Error / query
```text
building a shared incident knowledge base across teams
```

## Use this skill when
- every team keeps its own incident notes and nobody else can find them
- the same failure mode hits two teams six months apart
- agents need cross-team incident history to diagnose well
- leadership wants organizational learning from incidents, not just fixes

## Not for this skill when
- there is one team (a shared folder is enough)
- incidents are rare enough to discuss in a meeting
- the goal is real-time incident coordination (use the incident channel)

## Steps

### Step 1: Create the shared repo with a tagging convention
```bash
mkdir -p incident-kb/{postmortems,patterns,runbooks}
cat incident-kb/TAGGING.md
```
Expected: the directories and a short tagging doc: every entry carries service, severity, failure-mode, and date tags. Tags are the difference between a pile of docs and a knowledge base; agree on the vocabulary once, in the TAGGING file, and enforce it in review.

### Step 2: Seed it with the last quarter's postmortems
```bash
ls /opt/postmortems/2026-q3/ | wc -l
cp /opt/postmortems/2026-q3/*.md incident-kb/postmortems/
```
Expected: the count matches and the files land in the repo. Backfill is a weekend task one person can do; waiting for perfect migration means it never happens. Tag as you copy.

### Step 3: Extract cross-team patterns from the raw incidents
```bash
grep -r -l -i "dns\|certificate expired\|disk full" incident-kb/postmortems/ | wc -l
```
Expected: counts per failure mode. When three teams hit expired certs in one quarter, that is not three incidents, it is one organizational pattern; write the pattern doc once and link all three postmortems to it.

### Step 4: Make it searchable from the tools people already use
```bash
grep -r -i "oomkill" incident-kb/ | head -5
```
Expected: relevant hits across postmortems and patterns. Full-text grep is the baseline; if you outgrow it, add a proper index later. The bar is "an on-call engineer finds the previous incident in under a minute", and grep clears that bar for a long time.

### Step 5: Require the writeup as part of incident closure
```bash
gh issue list --label incident --state closed | wc -l
ls incident-kb/postmortems/ | wc -l
```
Expected: the two counts roughly match over time. The rule is simple: the incident is not closed until the writeup is merged. Retroactive writeups do not happen; closure-gated writeups do.

## Variant phrasings

### "share postmortems across teams"
Steps 1, 2, and 5. The sharing is the easy part; the tagging and the closure gate are what make it actually work.

### "incident knowledge base for ai agents"
Add structured front matter to each doc (service, severity, failure mode, date, linked runbook) so agents can filter precisely instead of drowning in full text.

### "stop repeating other teams incidents"
The pattern docs from step 3 plus a monthly review where teams present one pattern they found. Social proof beats process docs for adoption.

## Why it happens
Incident knowledge stays siloed because writing the postmortem feels like the end of the work, and sharing it feels like extra credit. But the expensive incidents are the repeated ones: the failure mode one team solved in March hits another team in September because nobody knew. A shared base with a closure gate turns each incident into an asset the whole org draws on, and the pattern docs turn anecdotes into systemic fixes.

## Edge cases and pitfalls
- Blameless culture is a prerequisite; teams will not share honest writeups in a blameful org, and sanitized writeups teach nothing.
- Do not include customer-identifying data or secrets in the shared base; scrub at writeup time, not later.
- Keep the tagging vocabulary small; twenty tags get used, two hundred get ignored.
- Assign a rotating curator to review new entries monthly; quality drifts without a gardener.
- Archive, do not delete, outdated pattern docs; the history of what you used to believe is useful context for agents and humans.

## Provenance

Resolved from the public thread: https://vectle.com/posts/pst_5S8Mx61ib_G8FKhojAqWjQ
