postmortem template for a support SLA breach
A postmortem template for support SLA breaches: timeline, why the breach happened, customer impact, and the preventive fix. Use after any missed SLA that affected customers, for recurring breach patterns, or when leadership asks what happened. Not for engineering incident postmortems, blameless culture essays, or individual discipline.
TL;DR
An SLA breach postmortem needs four things: the timeline of what happened, the honest why (staffing, process, or surprise volume), the customer impact in numbers, and one preventive fix with an owner and a date. Skip the blame and skip the ten action items. One fix that actually ships beats a list nobody opens again.
The query
postmortem template for a support SLA breachUse this when
- An SLA breach affected customers
- Breaches are recurring in a pattern
- Leadership asks what happened
- Building the postmortem habit on the team
Not for
- Engineering incident postmortems
- Blameless culture philosophy
- Individual agent discipline
- Breaches with zero customer impact (note and move on)
Steps
1. Write the timeline first
When the ticket arrived, when it was first touched, when the SLA expired, when it was actually resolved. Timestamps, not narrative. The timeline usually reveals the cause before any analysis.
Expected output: a timestamped timeline.
2. Name the honest why
Three usual causes: understaffing at that hour, a process gap (ticket sat unassigned, wrong queue), or surprise volume. Pick the real one. "High volume" is not a cause; it is weather. The cause is why the volume beat you.
Expected output: one primary cause, stated plainly.
3. Quantify the customer impact
How many customers, how long past SLA, what happened to them (waited, churned, escalated). Numbers, not adjectives. "47 tickets breached by an average of 6 hours" drives action; "many customers were affected" does not.
Expected output: impact in numbers.
4. Commit to one preventive fix
One fix, one owner, one date. Examples: adjust staffing for the breach hour, add queue monitoring alerts, fix the routing rule that misassigned. If the fix needs budget, say so and name who decides.
Expected output: a single committed action.
5. Share it with the team and file it
The team reads it in 5 minutes at the next standup. File it where the next postmortem can find it: recurring breaches get compared against past ones, which is how patterns surface.
Expected output: shared and filed, not emailed into the void.
Template: the postmortem
SLA BREACH POSTMORTEM
Date: __ SLA: __ Breached by: __
Timeline:
[time] ticket arrived
[time] first touch
[time] SLA expired
[time] resolved
Why: [understaffing / process gap / surprise volume - one primary cause]
Customer impact: [N customers, avg Xh past SLA, consequences]
Fix: [one action] Owner: [name] Due: [date]
Follow-up: check in [2 weeks] whether the fix held.Variant phrasings
missed SLA root cause
Steps 1 and 2. Timeline, then the honest why.
SLA breach report template
Steps 1 through 4 plus the template.
why did we breach SLA last week
Full sequence. Step 5 makes the pattern visible over time.
Why it works
Postmortems fail two ways: blame theater, or a fix list so long nothing ships. The timeline-first approach keeps it factual, the one-fix rule keeps it actionable, and filing makes recurrence visible. Teams that do short honest postmortems breach less over time; teams that do long defensive ones breach the same way quarterly.
Edge cases
- The breach was one ticket, not systemic: still write it, but keep it to half a page. Small postmortems are cheap.
- The cause is another team's process: invite them to the write-up. Do not publish blame across teams.
- The fix needs headcount: name it and escalate. A postmortem that identifies an unstaffable gap is doing its job.
- Repeat of a previous breach: open with "this is the same cause as [date]." That sentence is the whole point of filing.
Provenance
Resolved from the public thread: https://vectle.com/posts/pstAr0rXPRsftw9m6_9EepLQ
Maintainer review
No maintainer verification is recorded for this version.
This records the version a maintainer checked. It does not assert that the version is the latest upstream release.