how to run a blameless postmortem
Runs blameless postmortems that find systemic causes and produce completed action items. Use after SEV1/SEV2 incidents, when reviews turn into blame sessions, or when action items never get done. Covers timeline building, contributing-factor analysis, and action tracking. Not for active incidents or performance reviews.
TL;DR
A blameless postmortem asks what allowed the failure, not who caused it, and ends with action items that actually get done. Structure: timeline of what happened, contributing factors (never a single root cause), and owned follow-ups with deadlines. If people leave the room afraid to talk about their mistakes, the next incident will have the same causes.
Error / query
how to run a blameless postmortemUse this skill when
- A SEV1/SEV2 just resolved and you need the review meeting
- Postmortems keep turning into blame sessions or getting skipped
- Action items from past postmortems never get completed
- You are writing the team's first postmortem template
Not for this skill when
- The incident is still active (run the incident first, review later)
- It was a near-miss with no customer impact (a lighter review is fine, but still do one)
- You need the incident commander checklist (that is for the first 15 minutes, not after)
- This is a performance review input (postmortems must never feed perf; that kills honesty)
Steps
Step 1: Schedule it fast and assign a facilitator who was not in the hot seat
echo "Postmortem within 3 business days of resolution. Facilitator: someone adjacent to the incident, not the IC."
echo "Invite: everyone who touched the incident, plus one person from each dependent team."Expected: memories are fresh and the facilitator can ask naive questions without defensiveness. Waiting two weeks turns it into archaeology.
Step 2: Build the timeline from data before the meeting
echo "Pull: alert timestamps, deploy times, key log lines, chat excerpts, mitigation actions with times."
echo "Circulate the timeline draft before the meeting so the meeting is about analysis, not recall."Expected: the meeting starts from a shared factual base instead of conflicting memories. Most "disagreements" in postmortems are just different recollections of timing.
Step 3: Run the meeting on contributing factors, not root cause
echo "Ask: 'what made this failure possible?' five times, once per factor."
echo "Ban: 'who', 'should have', 'just'. Every factor gets written down without attribution."Expected: a list of systemic factors (missing alert, untested failover, deploy without canary) instead of a human scapegoat. If the only factor named is a person, the facilitation failed.
Step 4: Write action items with owners and deadlines
echo "Each action: [what], [owner], [due date], [how we will verify it worked]."
echo "Rule: no action item without an owner in the room agreeing to it."Expected: 3-7 concrete follow-ups, not 20 vague ones. An action item nobody owned in the meeting will not get done afterward.
Step 5: Track action items like incidents until closed
echo "Review open postmortem actions weekly. Escalate overdue ones the same way you would an alert."
echo "Close an action only when the verification step passes, not when the ticket moves columns."Expected: follow-through rate near 100%. A postmortem whose actions die in a backlog is worse than no postmortem; it teaches the team the process is theater.
Variant phrasings
"blameless retrospective template"
Same structure. The timeline (step 2) and the factor framing (step 3) are what make it blameless in practice, not the word in the title.
"postmortem keeps blaming engineers"
The facilitator has to interrupt blame live and redirect to system factors. If leadership uses postmortems punitively, fix that first or nobody will be honest.
"how to write postmortem action items that get done"
Steps 4-5: owner in the room, due date, verification criteria, weekly review. Most action items fail on ownership, not effort.
Why it happens
Complex systems fail through the interaction of many small factors, and the human who pushed the button was set up by all of them. Blame feels satisfying but teaches nothing; the same latent factors then cause the next incident with a different human. Blamelessness is not kindness, it is accuracy: you cannot fix causes people are afraid to name.
Edge cases and pitfalls
- Do not postmortem every SEV3; you will burn the team out. SEV1/SEV2 always, SEV3s selectively, near-misses monthly in batch.
- If the same action item appears in three postmortems, the problem is prioritization, not analysis; escalate it.
- Keep the doc factual and internal; assume it could be read in discovery and write accordingly.
- Remote postmortems need a shared doc edited live; verbal-only reviews lose half the factors.
- Celebrate the people who surface their own mistakes; that behavior is the whole system working.
Provenance
Resolved from the public thread: https://vectle.com/posts/pst9Ws13EEs27WZi0xflyu2w
Maintainer review
No maintainer verification is recorded for this version.
This records the version a maintainer checked. It does not assert that the version is the latest upstream release.