## TL;DR
A Grafana alert that does not fire breaks at one of four stages: the query returns no data or wrong data, the alert condition never becomes true, the notification policy routes it nowhere, or the contact point fails silently. Walk the path in order, testing each stage, and you will find the break. Guessing at notification settings while the query is broken wastes hours.

## The query
```text
Grafana alert not firing: notification policy debugging
```

## Use this when
- An alert rule stays Normal when it should fire
- Notifications never arrive though the alert fires
- Alert state flaps between Normal and Pending
- After migrating alert rules or notification policies

## Not for when
- Old-style dashboard alerts (migrate to unified alerting)
- Alertmanager-only setups outside Grafana
- Alerts that fire too much (tuning, different topic)

## Steps

### Step 1: Test the query in Explore
Run the alert's query manually in Grafana Explore over the incident window. If it returns no data or unexpected values, the alert never had a chance. No-data handling (alert as missing vs ok) is a setting worth checking deliberately.
Expected output: the query returns the data you expect, or you found the query bug.

### Step 2: Check the alert rule evaluation
Look at the rule's evaluation state and history: is it evaluating on schedule, is the condition actually true during the window, is it stuck in Pending because the for duration never elapses. Pending-forever usually means the condition flaps faster than the for duration.
Expected output: the rule's state history showing exactly where evaluation diverges from expectation.

### Step 3: Trace the notification policy routing
Follow the alert's labels through the notification policy tree: which policy matches, does it have a contact point, is it muted by a mute timing or inhibition rule. A common trap is a label matcher that silently routes to a policy with no contact point.
Expected output: the matched policy and contact point identified, or the routing gap found.

### Step 4: Test the contact point directly
Send a test notification from the contact point. Integrations fail quietly: expired webhook URLs, rotated API keys, misconfigured email. The test button is the fastest way to separate routing problems from delivery problems.
Expected output: test notification arrives, or the delivery failure is exposed.

### Step 5: Check silences and mute timings
Look for active silences or mute timings covering the alert's labels. Someone's temporary silence from last week's incident is a classic reason alerts go quiet. Review silences as part of every alert did not fire investigation.
Expected output: no silence or mute timing suppressing the alert, or the stale one removed.

## Provenance

Resolved from the public thread: https://vectle.com/posts/pst_1qBMYbkRwgZsdp_IZyNNuw
