how to use for clauses to stop alert flapping
Uses Prometheus for clauses to stop alert flapping. Use when alerts fire and resolve repeatedly, on-call gets paged for blips, or you need alerts that only fire on sustained conditions. Covers for duration tuning, combining with rate windows, and when for is not enough. Not for writing the alert expression itself, Alertmanager routing, or silencing.
TL;DR
A flapping alert is an alert without patience: add a for clause so the condition must hold continuously for a few minutes before firing. Start with 5 minutes for paging alerts and 15 for tickets, and make sure the underlying expression uses a rate or average window at least as long as the for duration. for does not fix a noisy expression; it gives a reasonable one time to settle.
Error / query
how to use for clauses to stop alert flappingUse this skill when
- Alerts fire and resolve in rapid cycles
- On-call is paged for transient blips
- You want alerts on sustained conditions only
- Tuning existing alerts that are too twitchy
Not for this skill when
- Writing the alert expression from scratch (PromQL guide)
- Routing or deduplicating in Alertmanager
- You want to silence an alert (silences, not tuning)
Steps
Step 1: Add a for clause to the flapping alert
- alert: HighErrorRate
expr: sum(rate(http_requests_total{status=~"5.."}[5m])) / sum(rate(http_requests_total[5m])) > 0.05
for: 5mExpected: the alert enters Pending when the condition first holds and only fires if it stays true for 5 continuous minutes. A 2-minute blip never pages; a real 10-minute incident does.
Step 2: Match the rate window to the for duration
# rule of thumb: rate/avg window >= for duration
expr: avg_over_time(cpu_usage[10m]) > 0.9
for: 10mExpected: the expression's own smoothing window is at least as long as the patience window. A 1-minute rate window with a 10-minute for still jitters, because the input is noisy even though the trigger is patient.
Step 3: Verify with promtool and the alert timeline
promtool check rules [rules-file]
# then in Prometheus UI: Alerts page shows Pending vs Firing historyExpected: the rule loads cleanly and the alert's history shows Pending periods that resolved without firing. Those suppressed pendings are the pages you just saved.
Step 4: Escalate the pattern for chronic flappers
If an alert still flaps with a sane for: the threshold is wrong
(not the timing). Widen to percentile-based thresholds, split
by label to find the noisy series, or alert on the burn rate
instead of the raw value.Expected: a decision. for handles timing noise; persistent flapping after tuning means the signal itself needs rethinking.
Variant phrasings
"prometheus alert keeps firing and resolving"
Steps 1-2. Add for, and lengthen the expression's window to match.
"alert flapping prometheus"
Same fix. Patience (for) plus smoothing (window) together.
Why it happens
Prometheus evaluates alert expressions on every cycle against instant or short-window data, so any threshold near a noisy signal's normal range fires on every excursion. The for clause requires the condition to persist, which filters out excursions shorter than real incidents while adding only minutes of detection delay.
Edge cases and pitfalls
for: 0m(or omitted) fires immediately; it is the default and the cause of most flapping.- Very long
for(30m+) delays real pages; reserve long durations for ticket-level alerts, not paging. forresets if the condition flickers false for even one evaluation; extremely spiky signals need the expression fixed, not longer patience.- Recording rules evaluated less often than the alert's evaluation can make
forbehave oddly; keep evaluation intervals aligned.
Provenance
Resolved from the public thread: https://vectle.com/posts/pst_saSqeTDE69aya2NQZECg
Maintainer review
No maintainer verification is recorded for this version.
This records the version a maintainer checked. It does not assert that the version is the latest upstream release.