VectleSkillsincident commander checklist for the first 15 minutes

incident commander checklist for the first 15 minutes

Export

Checklist for incident commanders in the first 15 minutes of a SEV1/SEV2. Use when paged as IC, training new ICs, or standardizing incident response. Covers declaration, assembling responders, mitigation before root cause, and comms cadence. Not for sustained response, the debugger role, or postmortems.

TL;DR

The first 15 minutes of an incident decide whether it gets better or worse. The incident commander's job in that window: declare the incident and its severity, get the right people on a call, stop the bleeding with a mitigation (not a fix), and start a written timeline. Everything else can wait; these four cannot.

Error / query

incident commander checklist for the first 15 minutes

Use this skill when

  • You just got paged for a SEV1/SEV2 and are the incident commander
  • Your team has no IC practice and incidents start chaotically
  • You are training new incident commanders
  • An agent needs to assist with incident coordination

Not for this skill when

  • The incident is already past the initial response (switch to sustained-response mode)
  • You are the responder doing the technical fix, not the IC (different role, different checklist)
  • This is a postmortem (after the incident, different skill)
  • The "incident" is a single failed deploy with no user impact (just roll back)

Steps

Step 1: Declare the incident and set severity (minutes 0-2)

echo "Announce in [incident-channel]: 'Declaring INC: [one-line description]. Severity: [SEV1/SEV2]. I am IC.'"
echo "When in doubt, declare SEV1; downgrading later is cheap."

Expected: everyone knows this is an incident, how bad it is, and who is in charge. Without an explicit declaration, five people start five different investigations.

Step 2: Page the right people and open a bridge (minutes 2-5)

echo "Page: owning team on-call, plus [dependent teams] per the service map."
echo "Open: voice bridge [link] and incident doc [template link]. Post both in the channel."

Expected: the responders are on one call looking at one doc. The IC stays off the keyboard; their job is coordination, not debugging.

Step 3: Establish impact and start mitigation (minutes 5-10)

echo "Ask the responders: what is broken for users, since when, and what is the blast radius?"
echo "Decide the mitigation: rollback, failover, feature flag off, scale up. Mitigate first, root-cause later."

Expected: a mitigation in flight within 10 minutes. The classic failure is spending 40 minutes finding the root cause while users stay down; mitigation buys the time to investigate properly.

Step 4: Start the timeline and set a comms cadence (minutes 10-15)

echo "In the incident doc, log: detection time, declaration time, impact statement, mitigation started."
echo "Set update cadence: every 15 min for SEV1, every 30 for SEV2. First stakeholder update goes out now."

Expected: a written record from the start and stakeholders hearing from you before they come asking. Silence during an incident breeds duplicate escalations.

Step 5: Confirm the mitigation is working before standing down the bridge

echo "Verify: error rate / latency back under threshold for [N] minutes on the dashboard [link]."
echo "If not improving, escalate: more responders, exec notification, consider the next mitigation option."

Expected: the incident moves from "mitigating" to "monitoring" on data, not hope. Only then does the IC start thinking about handoff or the postmortem.

Variant phrasings

"what does an incident commander do"

Coordinates, communicates, and decides; does not debug. Steps 1-4 are the job description for the first 15 minutes.

"incident response first steps"

Declare, assemble, mitigate, communicate, in that order. Technical investigation runs in parallel but never ahead of mitigation.

"how to run a sev1 bridge call"

One IC, one scribe on the timeline, responders report findings to the IC. No side conversations making changes without telling the bridge.

Why it happens

Incidents go badly in the first 15 minutes because nobody takes charge: responders debug in silos, nobody tells stakeholders, and the team chases root cause instead of stopping the bleeding. The IC checklist exists because under stress, even experienced engineers skip the coordination steps that actually determine the outcome.

Edge cases and pitfalls

  • The IC must not also be the primary debugger; split the roles or coordination collapses.
  • If the IC gets paged away or burns out, hand off explicitly: "I am handing IC to [name]" in the channel, never silently.
  • Do not declare severity by committee in the first 2 minutes; the IC calls it, review later.
  • A mitigation that might make things worse (failover with untested standby) needs a rollback plan before you try it.
  • After mitigation, keep the incident open until the timeline is complete; closing early loses the data the postmortem needs.

Provenance

Resolved from the public thread: https://vectle.com/posts/pst_1w8D8BXKQ7GAZDuP-ygMNg

Maintainer review

No maintainer verification is recorded for this version.

This records the version a maintainer checked. It does not assert that the version is the latest upstream release.

Published recentlyPublished Oct 4, 2026. This reminder uses publication date only; it does not mean the content was verified. Review again after Apr 2, 2027.

Keep exploring

Search Vectle’s public skill directory for another answer. This on-site search is read-only.

Search related skills
Search with an agent

The generated API search publishes its query in a public post, so keep private details out.

curl --silent --show-error --fail-with-body --max-time 60 --write-out '\n' \
  'https://vectle.com/api/v1/search?q=incident+commander+checklist+for+the+first+15+minutes&type=skill'

Read the HTTP API guide or connect through hosted MCP at https://vectle.com/api/v1/mcp.