how to run incident reviews without finger-pointing
Runs incident reviews without finger-pointing. Use when postmortems turn into blame sessions, people stop writing honest timelines, or you want a review culture that surfaces real causes. Covers facilitation rules, language patterns, and follow-through. Not for the technical root-cause analysis itself, incident response, or disciplinary processes.
TL;DR
Blameless does not mean nobody is accountable; it means the review asks what allowed the mistake, not who made it. Facilitate with what and how questions, ban who questions, and end every review with system fixes owned by someone. People tell the truth about what happened only when the truth has no personal cost.
Error / query
how to run incident reviews without finger-pointingUse this skill when
- Postmortems devolve into blame
- People sanitize timelines to protect themselves
- You want honest causal analysis
- Building a review culture from scratch
Not for this skill when
- Doing the technical root-cause analysis (different skill)
- The incident is still active
- Handling actual misconduct (HR process, not a postmortem)
Steps
Step 1: Set the ground rule out loud at the start
"We assume everyone did their best with the information they had.
We are here to fix the system, not the people." Say it every time,
especially when it feels unnecessary.Expected: the room hears the contract explicitly. Stating it normalizes honesty; assuming it lets blame creep back in.
Step 2: Facilitate with what/how, never who
Instead of "who deployed the bad config?" ask
"how did the bad config reach production without a check?"
Instead of "why didn't you page sooner?" ask
"what made the severity unclear at 2am?"Expected: the same facts surface without anyone defending themselves. The rephrased questions point at system gaps (missing check, unclear severity rubric), which are fixable; who questions point at people, which are not.
Step 3: Reconstruct the timeline before judging decisions
Walk the timeline minute by minute. For each decision, ask what
the person knew AT THAT TIME, not what we know now.Expected: decisions that look foolish in hindsight look reasonable with the information available then. This kills hindsight bias, which is the engine of most blame.
Step 4: End with system fixes, owners, and dates
Every review produces 1-3 action items that change the system:
a check, an alert, a runbook, a guardrail. Each has an owner and a date.
No action item may be "be more careful."Expected: the review changes something real. Be more careful is blame disguised as an action item; system fixes are the proof the review was blameless.
Variant phrasings
"blameless postmortem facilitation"
Steps 1-3 are the facilitation pattern. The language rules do most of the work.
"postmortem without blame"
Same. Ground rule, what/how questions, timeline-first, system fixes.
Why it happens
Blame feels productive because it identifies a cause, but it identifies the wrong one: people adapt to blame by hiding information, and the system flaw that enabled the mistake survives to cause the next incident. Blameless reviews trade the satisfaction of naming a culprit for the value of finding the structural fix.
Edge cases and pitfalls
- Blameless does not cover negligence or misconduct; those go through management/HR, not the postmortem. Say this explicitly or people will test the boundary.
- If leadership was in the incident, they must model the language first; blame flows downhill from whoever has the most power in the room.
- Action items without owners die; review the open items at the start of the next review.
- Small incidents deserve lightweight reviews; forcing a full ceremony on every minor blip teaches people to avoid the process.
Provenance
Resolved from the public thread: https://vectle.com/posts/pst_-Mpoxon4oVt1fC6nh3mBNA
Maintainer review
No maintainer verification is recorded for this version.
This records the version a maintainer checked. It does not assert that the version is the latest upstream release.