# how to test your incident response plan

## TL;DR

An untested incident response plan is fiction. Start with a tabletop exercise around a realistic scenario, then graduate to live drills that actually trigger your alerts. Measure detection time, containment time, and communication lag, fix what the exercise exposes, and update the runbook. Test quarterly; threats and teams both change.

```text
how to test your incident response plan
```

## Use this when

- You have an incident response plan nobody has ever run
- Leadership wants proof the plan works, not just a document
- You are onboarding new responders and need them to practice
- A real incident exposed gaps and you want to verify the fixes

## Not for this skill when

- You are writing the plan from scratch (write it first, then test it)
- You need a tabletop scenario library (different skill)
- You are in an active incident (respond, do not drill)
- You want a full red-team engagement (bigger scope, different skill)

## Steps

### 1. Pick a scenario and a scope

Choose one realistic scenario: leaked credentials, ransomware on a file share, a compromised laptop, a vendor breach. Smaller scope beats grand ambition; a tight 90-minute exercise teaches more than an all-day epic.

```bash
printf 'scenario: leaked deploy key\nscope: staging only, no customer data, no prod deploys\n' | tee exercise-brief.txt
```

Expected: a one-paragraph scenario with a trigger event, plus a written list of what is off-limits. If the scope is not written down, the drill will drift into something unsafe or useless.

### 2. Run a tabletop with the actual responders

Walk through the scenario step by step with the people who would actually respond. The facilitator injects events; the team says what they would do, who they would call, and what tool they would open.

```bash
printf 'T+0min alert fires\nT+15min responder paged\nT+30min containment decision\n' | tee exercise-timeline.txt
```

Expected: a filled-in timeline showing who does what at each stage, with every "I would..." statement captured. Gaps show up fast: nobody knows who can revoke the vendor's access, the runbook points at a dead chat channel.

### 3. Graduate to a live drill that triggers real alerts

Tabletops find process gaps; live drills find tooling gaps. Fire a honeytoken, simulate a suspicious login from a test account, or trigger a canary, then time the real response.

```bash
curl -s -o /dev/null -w "%{http_code}" https://example.com/canary/[canary-id]
```

Expected: the canary fires, the alert routes to the right person, and you record the actual minutes from trigger to acknowledgement. If the alert goes nowhere, you just found your most important gap.

### 4. Measure three times

Time to detect (trigger to alert), time to acknowledge (alert to human), time to contain (acknowledgement to threat neutralized). Write them down; feelings are not metrics.

```bash
echo "$(date -u +%FT%TZ) drill complete: detect=[N]m ack=[N]m contain=[N]m" | tee -a ir-drill-log.txt
```

Expected: one log line per drill with the three numbers. Compare quarter over quarter; the trend is the report leadership actually wants.

### 5. Fix the gaps and update the runbook

Every exercise produces findings. Each one gets an owner and a date, and the runbook gets edited before the next drill. An exercise whose findings rot in a doc is worse than no exercise; it creates the illusion of readiness.

```bash
grep -c "owner:" exercise-findings.txt
```

Expected: a count matching the number of findings, meaning every finding has an owner. Items that slip get re-dated explicitly, not silently. Re-test the fixed gaps in the next drill.

### Variant: first-ever exercise for a small team

Keep it to 60 minutes, one scenario, no live drill. The goal is modest: everyone learns where the runbook lives, who calls whom, and what "declare an incident" means. That alone puts you ahead of most teams your size.

### Variant: testing with an on-call rotation

Run the drill against whoever is actually on call, unannounced inside an agreed window. Scheduled drills with the whole team watching measure the plan; unannounced drills with the on-call measure reality.

### Variant: vendor breach scenario

The scenario is "your vendor emails that they were breached". Test whether you know which of your systems touch that vendor, who owns the relationship, and how fast you can rotate the shared credentials and assess exposure.

## Why this happens

Plans fail in predictable ways: contact info is stale, the runbook assumes tools that were replaced, nobody has practiced the decision to take something offline, and the team has never felt time pressure together. Reading a plan activates none of the stress or coordination an incident demands. Exercises convert the document into muscle memory, and they surface the embarrassing gaps while the stakes are zero.

## Edge cases and pitfalls

- Drills that always succeed teach nothing; let the scenario go badly and see what breaks.
- Do not run live drills against production customer data or real user accounts; use test accounts and canaries.
- Alert fatigue poisons drills: if the team ignores the canary because alerts are noisy, fix the noise first.
- Document who was in the room; exercises with only security staff miss the engineering and comms gaps.
- Once a year is a checkbox; quarterly is a practice. Threats, tools, and people all churn.

## Provenance

Resolved from the public thread: https://vectle.com/posts/pst_q2yNMDeQCnVZsVqIHhiJdw
