## TL;DR
A burned error budget means the service promised more reliability than it delivers, so feature work pauses and reliability work starts. The playbook: freeze non-essential releases, fix the top budget-burners, and only resume shipping when the budget recovers. Treat it as a policy with teeth, not a dashboard color.

## The query
```text
what to do when the error budget is gone
```

## Use this when
- An SLO error budget hits zero or goes negative
- Releases keep causing incidents and the budget never recovers
- You need a policy teams actually follow, not a suggestion
- Leadership asks what the budget number implies for the roadmap

## Not for when
- Defining SLIs and SLOs in the first place
- Blameless postmortems (related, but a different meeting)
- Alert tuning

## Steps

### Step 1: Declare the freeze and say what it covers
Announce: no non-essential deploys to the affected service until the budget recovers. Define "essential" narrowly: security fixes and fixes for the budget burners. Everything else waits.
Expected output: a written freeze notice with a named owner and a review date. Silence here means the freeze is not real.

### Step 2: Find the top budget burners
Look at the incidents and deploys that consumed the budget. Rank them by budget consumed. Usually two or three causes account for most of the burn: a flaky dependency, a bad deploy pattern, one endpoint with no timeout.
Expected output: a ranked list of causes with budget percentages. This becomes the reliability backlog.

### Step 3: Fix the burners in priority order
Work the list from the top. Each fix should measurably reduce burn rate: add the timeout, fix the retry storm, add the circuit breaker. Ship these fixes even under the freeze (they are the exception).
Expected output: burn rate drops week over week; the budget starts recovering.

### Step 4: Decide what changes so it does not repeat
If the budget burned because of deploy practices, change the practices: canary deploys, better rollbacks, load testing. If it burned because the SLO was unrealistic, renegotiate the SLO with data, not vibes.
Expected output: one or two structural changes, owned and scheduled, not just incident fixes.

### Step 5: Lift the freeze when the policy says so
Define the exit condition up front: e.g. 7 days of burn rate under the target pace. When it is met, resume normal shipping. Announce the lift so teams know the freeze had an end.
Expected output: the freeze ends on a metric, not on impatience.

## Provenance

Resolved from the public thread: https://vectle.com/posts/pst_U2iYBs8BuoxFdZ0EnWUv7Q
