how to write a PromQL alert that doesn't flap
Shows how to write a PromQL alert that does not flap: alert on rates over windows, aggregate away per-pod noise, and require persistence with for. Use when alerts fire and resolve repeatedly. Validated with promtool; does not cover Alertmanager routing.
TL;DR
Use for to require the condition to persist, aggregate away per-pod noise with sum or avg, and alert on rates over sensible windows instead of instant values. Flapping alerts are almost always instant-value comparisons on noisy signals. A 5-minute for plus a rate window fixes most of them.
Error / query
how to write a PromQL alert that doesn't flapYour alert fires, resolves, fires again, and the team has stopped believing it. You need the alert to mean something.
Use this skill when
- an alert keeps firing and resolving on its own
- the team ignores an alert because it cried wolf too often
- youre writing a new alert and want it stable from day one
- youre auditing alerts that page too often
Not for this skill when
- the alert never fires at all (a threshold problem, not flapping)
- you need Alertmanager routing or silencing config (delivery, not the expression)
- the underlying metric has no data (see the "no data" skill)
Steps
1. Rewrite the alert around a rate, not an instant value
Instant values bounce; rates over a window dont. Aggregate across pods so one noisy replica cant flap the alert alone.
cat > /tmp/alert-noflap.yaml <<'EOF'
- alert: HighErrorRate
expr: sum(rate(http_requests_total{status=~"5.."}[5m])) / sum(rate(http_requests_total[5m])) > 0.05
for: 10m
labels:
severity: page
annotations:
summary: "error rate above 5 percent for 10 minutes"
EOF
cat /tmp/alert-noflap.yamlExpected: the alert only fires when the error rate is genuinely elevated for 10 minutes. One bad scrape cant trip it.
2. Validate the expression before deploying it
Catch syntax errors locally instead of discovering them in production at 2am.
promtool check rules /tmp/alert-noflap.yamlExpected: SUCCESS. promtool catches syntax and structural errors before the rule ever reaches the server.
3. Backtest how often it would have fired
Run the expression in the Prometheus UI over the last 7 days and count the crossings. If it still flaps historically, lengthen for or widen the rate window before shipping.
echo "Paste the expr into the Prometheus graph UI with a 7-day range."
echo "Count the fire/resolve cycles. More than 2 per week means lengthen 'for:' or widen the rate window."Expected: at most a couple of crossings per week on historical data. The alert you ship is the alert you backtested.
Variant phrasings
prometheus alert keeps firing and resolving
Same fix: the for clause plus rate-based expression in step 1 is the direct cure for fire-resolve cycling.
how to stop alert flapping
Same fix: persistence windows and aggregation. Flapping is a signal-processing problem, and these are the filters.
promql for clause best practices
Same fix: set for to at least 2-3x your scrape interval, longer for noisy signals. Backtest per step 3.
Why it happens
Instant vectors bounce around with every scrape: GC pauses, deploy blips, one slow request. Comparing an instant value to a threshold turns that normal jitter into pages. Rates smooth the jitter, aggregation removes single-replica noise, and for demands persistence. Flapping isnt the signal lying; its the alert asking the wrong question of the signal.
Edge cases and pitfalls
- Batch jobs with spiky traffic: use longer rate windows that cover the spike pattern, or the alert learns the wrong baseline.
forlonger than the incident: a 30-minuteforon a 10-minute outage never pages. Match the window to the response target.- Alerts on
absent(): pair them with a deadmans-switch test so you know the alerting pipeline itself is alive. - Recording rule feeding the alert is stale: check the rule evaluation interval. A 5-minute rule cant feed a 1-minute alert.
Provenance
Resolved from the public thread: https://vectle.com/posts/pst_G1GVyJsVmgA9lNuTlcldAw
Maintainer review
No maintainer verification is recorded for this version.
This records the version a maintainer checked. It does not assert that the version is the latest upstream release.