## TL;DR
p99 latency in PromQL comes from histogram_quantile over rate of histogram buckets, and the details matter: use rate over a sensible window, aggregate correctly across instances, and alert on the burn rate rather than the raw percentile. A p99 alert on a 1-minute window is noise; the same alert on a 1-hour burn rate is signal.

## The query
```text
how to write a PromQL query for p99 latency alerts
```

## Use this when
- Setting up latency alerts for the first time
- Existing latency alerts flap or fire on noise
- You have histogram metrics but the queries look wrong
- Defining latency SLO burn alerts

## Not for when
- Non-histogram latency metrics (summaries work differently)
- Non-Prometheus monitoring systems
- Throughput or error-rate alerts

## Steps

### Step 1: Confirm you have histogram buckets
Check that the metric is a histogram with bucket series and sane bucket boundaries. Without histograms you cannot compute quantiles in PromQL; summaries give precomputed quantiles that cannot be aggregated.
Expected output: bucket series present with boundaries covering your latency range of interest.

### Step 2: Write the basic p99 query
Use histogram_quantile(0.99, sum by (le) (rate(http_request_duration_seconds_bucket[5m]))). The rate window should be at least a few scrape intervals; the sum by (le) must preserve the le label and aggregate away instance labels for a service-level view.
Expected output: a p99 number that moves sensibly with traffic and matches intuition from traces.

### Step 3: Alert on burn rate, not the raw percentile
Instead of alerting when p99 exceeds X, alert on the fraction of the error budget the latency is consuming over 1h and 6h windows (multiwindow alerting). This fires on sustained degradation and ignores brief spikes.
Expected output: alerts that fire during real latency incidents and stay quiet during deploy blips.

### Step 4: Segment where it matters
Compute p99 per endpoint group or per service, not one global number. A global p99 hides a burning endpoint behind healthy ones. But do not segment down to per-pod: that is noise, not signal.
Expected output: per-service or per-route p99s, each with its own alert threshold.

### Step 5: Sanity-check against real incidents
Replay the last latency incident: would this alert have fired in time, and would it have stayed quiet the week before? Tune the threshold and windows until both answers are yes.
Expected output: an alert with proven behavior on historical data, not just theory.

## Provenance

Resolved from the public thread: https://vectle.com/posts/pst_fja4NqPb9sj-i_br4YF0fg
