alert fatigue: how to tune noisy alerts
Explains how to tune noisy alerts and beat alert fatigue: rank alerts by firings without action, demote or delete the worst offenders, and add persistence windows. Use when on-call is drowning in pages nobody acts on. Does not cover writing new alerts from scratch.
TL;DR
Kill or demote every alert that fires more often than it leads to action. The fix for fatigue is fewer, better alerts, not a better on-call attitude. An alert nobody acts on trains everyone to ignore all of them. Start by measuring which alerts fire without action, then delete or downgrade the top offenders.
Error / query
alert fatigue: how to tune noisy alertsThe pager goes off all night and nobody learns anything from it. People have started ignoring alerts, including the real ones.
Use this skill when
- on-call is drowning in pages nobody acts on
- alerts fire and resolve on their own repeatedly
- the team has started muting or ignoring the pager
- youre auditing an alert ruleset that grew without pruning
Not for this skill when
- you need to write new alerts from scratch (a different skill)
- the problem is missing alerts, not noisy ones (coverage gap, opposite fix)
- you need Alertmanager routing config (delivery mechanics, not tuning)
Steps
1. Find your noisiest alerts
Count alert firings over the last 7 days. The ones at the top of the list are your noise.
curl -s --get "http://YOUR-prometheus-host/api/v1/query_range" \
--data-urlencode "query=ALERTS" \
--data-urlencode "start=$(date -u -d '7 days ago' +%Y-%m-%dT%H:%M:%SZ)" \
--data-urlencode "end=$(date -u +%Y-%m-%dT%H:%M:%SZ)" \
--data-urlencode "step=1h" -o /tmp/alerts.json
python3 -c "
import json
d = json.load(open('/tmp/alerts.json'))
series = d['data']['result']
ranked = sorted(series, key [your value] s: len(s['values']), reverse=True)
for s in ranked[:10]:
print(len(s['values']), s['metric'].get('alertname'), s['metric'].get('severity'))
"Expected: the top 10 alerts by firing frequency. These are the candidates for demotion or deletion.
2. Demote or delete the top offenders
For each noisy alert, pick one: fix the threshold, demote page to ticket, or delete it. "Keep and ignore" is not an option.
grep -rn "severity: page" ./prometheus-rules/ | head -20Expected: a list of paging alerts to review one by one. Each gets a fix, a demotion, or a deletion, decided this week, not someday.
3. Add a persistence window so transient blips dont page
Most flapping dies with a for clause. The alert should only page when the condition persists, not on one bad scrape.
cat > /tmp/alert-snippet.yaml <<'EOF'
# before: fires on any single scrape miss
# after: must persist 5 minutes before paging
- alert: ServiceDown
expr: up == 0
for: 5m
labels:
severity: page
EOF
cat /tmp/alert-snippet.yamlExpected: alerts only page when the condition holds for the full window. Single-scrape blips stop waking people up.
Variant phrasings
how to reduce alert noise
Same fix: rank by noise (step 1), then demote, fix, or delete (step 2). Noise reduction is subtraction, not better thresholds alone.
too many prometheus alerts
Same fix: the method above is Prometheus-shaped but the principle is universal. Count firings, cut the ones with no action.
how to tune alerting thresholds
Same fix: tune only the alerts that survive step 2. Tuning an alert nobody acts on is polishing noise.
Why it happens
Alerts are easy to add and nobody owns removing them, so every incident adds alerts and none ever get deleted. Each ignored alert teaches the team that alerts dont matter, which is exactly backwards from what paging is for. Fatigue is a rational response to noise; the fix is removing the noise, not lecturing the responders.
Edge cases and pitfalls
- An alert nobody understands: delete it. If nobody knows what action it wants, it cant be actionable.
- Compliance-required alerts: keep them, but route to a ticket queue, not a pager. Compliance doesnt require waking humans.
- Seasonal traffic patterns: dont tune thresholds to the quiet season; use windows that survive the busy one.
- Alert fatigue already severe: declare an alert amnesty week. Review every paging alert and justify its existence out loud.
Provenance
Resolved from the public thread: https://vectle.com/posts/pst_ttYzqQH2iC2Fgh5hSartaQ
Maintainer review
No maintainer verification is recorded for this version.
This records the version a maintainer checked. It does not assert that the version is the latest upstream release.