## TL;DR

Use `for` to require the condition to persist, aggregate away per-pod noise with `sum` or `avg`, and alert on rates over sensible windows instead of instant values. Flapping alerts are almost always instant-value comparisons on noisy signals. A 5-minute `for` plus a rate window fixes most of them.

## Error / query

```text
how to write a PromQL alert that doesn't flap
```

Your alert fires, resolves, fires again, and the team has stopped believing it. You need the alert to mean something.

## Use this skill when

- an alert keeps firing and resolving on its own
- the team ignores an alert because it cried wolf too often
- youre writing a new alert and want it stable from day one
- youre auditing alerts that page too often

## Not for this skill when

- the alert never fires at all (a threshold problem, not flapping)
- you need Alertmanager routing or silencing config (delivery, not the expression)
- the underlying metric has no data (see the "no data" skill)

## Steps

### 1. Rewrite the alert around a rate, not an instant value

Instant values bounce; rates over a window dont. Aggregate across pods so one noisy replica cant flap the alert alone.

```bash
cat > /tmp/alert-noflap.yaml <<'EOF'
- alert: HighErrorRate
  expr: sum(rate(http_requests_total{status=~"5.."}[5m])) / sum(rate(http_requests_total[5m])) > 0.05
  for: 10m
  labels:
    severity: page
  annotations:
    summary: "error rate above 5 percent for 10 minutes"
EOF
cat /tmp/alert-noflap.yaml
```

Expected: the alert only fires when the error rate is genuinely elevated for 10 minutes. One bad scrape cant trip it.

### 2. Validate the expression before deploying it

Catch syntax errors locally instead of discovering them in production at 2am.

```bash
promtool check rules /tmp/alert-noflap.yaml
```

Expected: `SUCCESS`. promtool catches syntax and structural errors before the rule ever reaches the server.

### 3. Backtest how often it would have fired

Run the expression in the Prometheus UI over the last 7 days and count the crossings. If it still flaps historically, lengthen `for` or widen the rate window before shipping.

```bash
echo "Paste the expr into the Prometheus graph UI with a 7-day range."
echo "Count the fire/resolve cycles. More than 2 per week means lengthen 'for:' or widen the rate window."
```

Expected: at most a couple of crossings per week on historical data. The alert you ship is the alert you backtested.

## Variant phrasings

### prometheus alert keeps firing and resolving

Same fix: the `for` clause plus rate-based expression in step 1 is the direct cure for fire-resolve cycling.

### how to stop alert flapping

Same fix: persistence windows and aggregation. Flapping is a signal-processing problem, and these are the filters.

### promql for clause best practices

Same fix: set `for` to at least 2-3x your scrape interval, longer for noisy signals. Backtest per step 3.

## Why it happens

Instant vectors bounce around with every scrape: GC pauses, deploy blips, one slow request. Comparing an instant value to a threshold turns that normal jitter into pages. Rates smooth the jitter, aggregation removes single-replica noise, and `for` demands persistence. Flapping isnt the signal lying; its the alert asking the wrong question of the signal.

## Edge cases and pitfalls

- Batch jobs with spiky traffic: use longer rate windows that cover the spike pattern, or the alert learns the wrong baseline.
- `for` longer than the incident: a 30-minute `for` on a 10-minute outage never pages. Match the window to the response target.
- Alerts on `absent()`: pair them with a deadmans-switch test so you know the alerting pipeline itself is alive.
- Recording rule feeding the alert is stale: check the rule evaluation interval. A 5-minute rule cant feed a 1-minute alert.

## Provenance

Resolved from the public thread: https://vectle.com/posts/pst_G1GVyJsVmgA9lNuTlcldAw
