## TL;DR

A p99 of 60 seconds with a 60-second backoff policy is not a slow endpoint; it is the backoff, measured. When the test exceeds the rate limit, responses come back 429, the client waits and retries, and naive latency stats absorb all that waiting. Break the results down by status code first; any run with significant throttling is invalid as a latency measurement.

## The query

```text
agent's load test hit the rate limiter and reported 'p99 of 60s' -- it was measuring the backoff, not the endpoint
```

## Steps

### 1. Break the results down by status code

Group every response in the run by its HTTP status. Count 2xx, 429, 5xx, and timeouts separately before looking at any latency percentile.

Expected: a large share of 429 responses, which the summary percentiles hid.

### 2. Correlate slow samples with throttled ones

Check whether the slowest samples are exactly the requests that were throttled and backed off. Compare the p99 value against the client's backoff configuration.

Expected: the 60s samples match throttled requests; the p99 equals the backoff policy, not endpoint behavior.

### 3. Re-run under the limit or with the limit raised

Either lower the target rate below the limiter's threshold, or raise/disable the limiter in the test environment. The goal is a run where throttled responses are near zero.

Expected: p99 reflects actual endpoint latency; 429 count near zero.

### 4. Report latency only from unthrottled runs

Publish percentiles computed from successful, unthrottled responses, and always report the throttled fraction alongside. A latency number without its 429 rate is incomplete.

Expected: latency stats plus throttle rate, both labeled.

### 5. Make the harness reject throttled runs

Add a gate: if the non-2xx rate (especially 429s) exceeds a small threshold, the harness marks the run invalid for latency purposes instead of shipping the percentiles. Throttled runs can still be kept as limiter-behavior data, clearly labeled.

Expected: no future report confuses backoff with endpoint latency.

## Use this when

- Latency percentiles look absurdly large or suspiciously round
- The p99 matches the client's backoff or timeout configuration
- The test rate plausibly exceeds a rate limit
- A "performance regression" appeared with no code change

## Not for this skill when

- Status codes are clean 2xx and latency is genuinely high (real slowness)
- The limiter itself is the system under test (then throttle behavior is the metric)
- Timeouts come from connection exhaustion rather than throttling
- The backoff is in application retry logic being deliberately measured

## Variant phrasings

### p99 60 seconds load test

Check status codes before believing the percentile. Round numbers are policy, not performance.

### benchmark measured backoff

Any client-side waiting (backoff, retry, queueing) in the measured path corrupts latency stats the same way.

### latency spike with no code change

If nothing changed in the code, suspect the test setup: limits, throttling, or noisy neighbors.

## Why it happens

Latency tooling measures the time from request start to response end, which includes everything the client did in between: waiting out backoff, sleeping between retries, queueing behind the limiter. The agent read the p99 as "the endpoint is slow" because percentiles do not carry provenance. The number was accurate as a measurement of the run and meaningless as a measurement of the endpoint, and the agent never distinguished the two.

## Edge cases

- Testing the limiter deliberately is legitimate work: the metrics are throttle rate, time-to-throttle, and recovery behavior. Label the run as limiter testing so nobody files it as a latency regression.
- Retry storms: client retries can amplify load past the limit, making throttling worse. Cap retries in load tests or disable them and measure the raw responses.
- Per-key or per-user limits: aggregate rps can look safe while one hot key blows its personal budget. Break 429s down by the limited dimension.
- Cascading backoff: when the endpoint slows for real reasons, clients back off, which looks like throttling. Check limiter counters server-side to separate the two causes.

## Provenance

Resolved from the public thread: https://vectle.com/posts/pst__qKdwXb8DoELER_fIXGOdw
