tabletop exercises for on-call teams
Runs tabletop exercises for on-call teams. Use when you want incident practice without touching production, testing whether runbooks and escalation actually work, or onboarding new on-call engineers. Covers scenario design, facilitation, and capturing gaps. Not for live game days, chaos experiments, or real incident response.
TL;DR
A tabletop exercise is a talk-through of a fake incident: the facilitator narrates the scenario (it is 2am, the primary database is unreachable), and the team says what they would do, step by step, without touching anything. It exposes gaps in runbooks, paging, and decision-making for the cost of a one-hour meeting. Run one per quarter per critical system.
Error / query
tabletop exercises for on-call teamsUse this skill when
- Practicing incident response without production risk
- Testing whether runbooks and escalation work
- Onboarding new on-call engineers
- Finding gaps before a real incident does
Not for this skill when
- You want live failure injection (game day / chaos)
- The incident is real (respond, do not exercise)
- Testing code (this is a people and process drill)
Steps
Step 1: Pick a scenario from your real risks
Good scenarios: primary database unreachable at 2am, deploy
breaks checkout during peak, provider region outage, on-call
primary unreachable (test the escalation), secrets rotation gone wrong.
Pick one per session; depth beats breadth.Expected: a scenario the team recognizes as plausible. Real past incidents (sanitized) make the best scenarios because the gaps are known to matter.
Step 2: Facilitate as a narrated walk-through
Facilitator: "Pager goes off: database unreachable. You are primary.
What is your first action?" Team answers; facilitator advances the
scenario based on their answers ("you check the dashboard; it shows...").
No laptops fixing things; talk through every step.Expected: the team verbalizes the response path: acknowledge, assess, escalate, mitigate, communicate. The facilitator injects complications when the team gets comfortable (the secondary does not answer; the runbook link is dead).
Step 3: Capture every gap and assumption
Scribe notes: runbook steps that do not exist, tools nobody can
access at 2am, decisions nobody is authorized to make, pages
that would go to the wrong person. These are the deliverables.Expected: a concrete gap list. The exercise is only as valuable as what gets written down; verbal we should fix that evaporates.
Step 4: Turn gaps into action items with owners
Each gap becomes a ticket: fix the runbook, fix the page routing,
grant the access, document the decision authority. Review at the
next exercise whether they got done.Expected: the next tabletop starts by checking last time's gaps. Closed gaps prove the exercises work; open ones prove where the process is stuck.
Variant phrasings
"incident response tabletop"
Steps 1-2. Scenario plus narrated walk-through, no production touched.
"on-call drill without outage"
Tabletop (this skill) for process; game days for live systems. Start here.
Why it happens
Most incident-response failures are not technical: the runbook is stale, the page went to someone on vacation, nobody knows who can authorize the failover. Tabletops surface exactly these failures in a safe hour, because talking through the response forces the team to confront each step instead of assuming it works.
Edge cases and pitfalls
- Do not let the exercise become a blame review of a past incident; it is a forward-looking drill, not a postmortem.
- Include the actual on-call roster, not just senior engineers; the exercise tests the real responders.
- Keep it to an hour; longer sessions lose focus and attendance drops next quarter.
- Vary scenarios; running the same database drill quarterly teaches the team to pass the drill, not to respond.
Provenance
Resolved from the public thread: https://vectle.com/posts/pst4rxr493iQubr7XfmBtNyQ