P1 ticket first-15-minutes response playbook
A minute-by-minute playbook for the first 15 minutes of a P1 ticket: confirm severity, page responders and name roles, post the first customer update, shield the lead from ticket noise, and set the clock for the next 15. Use when a P1 lands, when writing the support runbook, or when postmortems flag slow initial response. Not for full incident lifecycle, engineering root cause, or P2/P3 handling.
TL;DR
The first 15 minutes of a P1 are about containment and communication, not the fix. Confirm the impact, page the right people, post the first customer update, and protect the responder from ticket noise. Teams that follow a clock in the first quarter hour resolve faster than teams that jump straight to debugging.
The query
P1 ticket first-15-minutes response playbookUse this when
- A P1 or sev-1 ticket just landed
- You are writing the incident response runbook for support
- New leads freeze when the big red ticket arrives
- Postmortems keep saying "slow initial response"
Not for
- The full incident management lifecycle
- Engineering root-cause procedures
- P2/P3 ticket handling
- Blameless postmortem facilitation
Steps
1. Minutes 0-2: confirm it is really a P1
Check the blast radius: how many customers, is core function down, is there a workaround? A loud single customer is urgent but not a P1. Confirm or downgrade in two minutes, out loud to the team.
Expected output: a confirmed or downgraded severity, announced in the team channel.
2. Minutes 2-5: page the responders, name the roles
Page engineering on-call and name three roles: incident lead, comms owner, and scribe. Named roles prevent five people debugging and nobody talking to customers.
Expected output: three named roles posted in the incident channel.
3. Minutes 5-8: post the first customer update
A short public note on the ticket and status page: what is affected, who is affected, next update in 30 minutes. Customers who see this dont file duplicate tickets.
Expected output: first customer-facing update live.
4. Minutes 8-12: shield the responder
Route all incoming P1-related tickets to one queue view and assign someone to merge duplicates. The incident lead debugs; everyone else handles the noise. Unshielded leads drown in pings.
Expected output: duplicate tickets merging into the master ticket.
5. Minutes 12-15: set the clock for the next 15
Confirm the next update time, confirm the roles are filled, and confirm the customer impact statement is accurate. Then the lead goes heads-down on the fix with comms running on rails.
Expected output: next-update time posted, roles confirmed, impact statement current.
Ready-to-use template
P1 FIRST 15 MINUTES
0-2 min: Confirm severity. Customers affected: [count/scope].
Core function down? [yes/no]. Workaround? [yes/no].
2-5 min: Page on-call. Roles: Lead [name], Comms [name], Scribe [name].
5-8 min: First update out. "We are investigating [what] affecting [who].
Next update at [time]."
8-12 min: Shield the lead. Merge duplicates into [master ticket].
One queue view for all P1 traffic.
12-15 min: Confirm next update time, roles, impact statement.
Lead goes heads-down.Variant phrasings
sev-1 first response playbook for support
Same clock, different label. Steps 1 and 2 are where sev-1 responses usually fail.
what to do in the first 15 minutes of a major incident
Steps 3 and 4. Early customer communication plus shielding the responder is the whole game.
P1 incident response checklist for support teams
The template as-is. Print it, pin it, run it.
Why it happens
P1s go badly in the first 15 minutes because everyone does the heroic thing: jumps into debugging. Five heroes debugging means nobody confirmed the scope, nobody told the customers, and nobody is writing anything down. The clock forces the unglamorous work first, and the unglamorous work is what makes the fix land faster.
Edge cases
- False P1: downgrade fast and say so publicly. "We are downgrading this to P2, here is why" protects the P1 signal for next time.
- P1 at 3am with a skeleton crew: one person runs the clock and comms, engineering debugs. Shrink the roles, keep the updates.
- Two P1s at once: separate leads, separate channels, one comms owner across both. Never merge two incidents into one thread.
- The fix is found in minute 10: still post the update and still set the next check. "Fix deploying, monitoring" is a complete minute-12 update.
- Customer demands a call immediately: the comms owner takes it, not the lead. Protect the debugger at all costs.
Provenance
Resolved from the public thread: https://vectle.com/posts/pst_Yy5P14eeN0h6OkgS3RR9DQ
Maintainer review
No maintainer verification is recorded for this version.
This records the version a maintainer checked. It does not assert that the version is the latest upstream release.