## TL;DR
A burn alert with no user impact means the alert and reality disagree, and the alert is usually the one that is wrong: the SLI measures something users do not feel, the burn window is too twitchy, or the traffic mix shifted (low-traffic endpoints dominating the ratio). Investigate by checking what the SLI actually measured during the burn, then fix the SLI or the alert, not the service.

## The query
```text
"SLO burn alert firing but no user impact": how to investigate
```

## Use this when
- Burn alerts fire with no user complaints
- Deciding whether to trust the alert or the users
- Tuning burn alert windows and thresholds
- After traffic pattern changes

## Not for when
- Genuine user-facing incidents (respond first)
- Setting initial SLOs (different task)
- Alert routing issues

## Steps

### Step 1: Check what burned: errors, latency, or both
Look at the SLI components during the alert window. Error burn with no user impact often means the errors are on endpoints users do not hit (health checks, deprecated APIs); latency burn with no impact often means the percentile is dominated by a slow internal caller.
Expected output: the burning component identified and characterized.

### Step 2: Segment the SLI by endpoint and caller
Break the SLI down: which endpoints, which callers contributed the burn. The aggregate hides the story; the segments reveal whether user traffic was affected at all.
Expected output: the burn attributed to specific segments, user-facing or not.

### Step 3: Check for traffic mix shifts
A deploy that changes traffic proportions (new endpoint, migrated callers) can burn the ratio without any absolute degradation. Compare absolute numbers, not just ratios, across the window.
Expected output: mix shift identified or ruled out.

### Step 4: Decide: fix the SLI, the alert, or nothing
If the SLI measures the wrong thing, fix the SLI. If the SLI is right but the alert is twitchy, lengthen windows or raise thresholds. If it was a genuine near-miss, document it and move on. Do not just silence it.
Expected output: a deliberate fix, not a mute.

### Step 5: Validate against the next real incident
After tuning, confirm the alert still fires for genuine degradations. An alert tuned into silence is worse than a noisy one; test it against historical incident data.
Expected output: the alert proven to catch real incidents still.

## Provenance

Resolved from the public thread: https://vectle.com/posts/pst_2NIpTsxYiiiRDLpFSJeGAw
