# agent sampled only successful requests and concluded the endpoint was fast -- the timeouts were the whole story

## TL;DR
Re-sample including failed and timed-out spans, and lead with the timeout rate, not the latency of successes. Latency of successful requests tells you nothing about the requests that never completed, and those are usually the ones users complain about. The fix is to sample by span status and alert on timeout rate first.

## The misleading reading

```text
agent sampled only successful requests and concluded the endpoint was fast -- the timeouts were the whole story
```

## Steps

1. Count outcomes first, before any latency math. Query your trace store or logs for the endpoint grouped by status over the incident window:

```
# count by outcome, not by latency
SELECT status, COUNT(*) FROM spans
WHERE endpoint = '[endpoint]' AND hour = '[incident hour]'
GROUP BY status
```

   Expected: a large share of timeout, error, or cancelled spans sitting next to the 'fast' successes.

2. Check how the agent's sampler treats failures. Many samplers drop error spans by default, or the agent filtered status != ok when pulling data.

   Expected: you find the filter or the sampler config that excluded the failures.

3. Re-run the latency analysis on the full set, and report three numbers: success rate, timeout rate, and p99 latency INCLUDING timed-out spans at their timeout value.

   Expected: the 'fast endpoint' becomes an endpoint with a bad timeout rate; the p99 jumps because timeouts now count.

4. Find what the timeouts have in common: a downstream dependency, a specific parameter, a saturated pool. Correlate timeouts with dependency latency.

   Expected: one downstream call or resource whose latency matches the timeout pattern.

5. Fix the timeout cause (dependency, pool size, query), then set the sampler to keep a fixed share of error spans permanently and alert on timeout rate.

   Expected: future profiles show the timeout rate as the headline number, not a hidden one.

## Use this when

- A profiler agent reports a fast endpoint but users or error trackers report failures.
- Latency dashboards look healthy during an incident.
- The analysis mentions only p50/p95/p99 of successful requests.

## Not for this skill when

- The timeout rate is genuinely near zero (then the successes really are the story).
- Failures are client-side aborts (user navigated away), not server timeouts; those need different handling.
- You are profiling a batch job where 'timeout' is not a meaningful outcome category.

## Variant phrasings

### p99 looks fine but error rate is spiking
The latency math excluded the errors.

### agent says endpoint is fast, users say it hangs
Hangs are timeouts the agent filtered out.

### traces show 40ms but support tickets say timeouts
The trace sample kept only the completions.

## Why it happens
Timeouts do not produce normal latency samples: the span either never closes, gets recorded with a synthetic duration, or is dropped by a sampler that only keeps completed spans. Analysts and agents then compute percentiles over successes only, which is conditioning on the outcome you are trying to detect. The endpoint looks fast precisely because the slow requests were removed from the data.

## Edge cases

- Retries turn one user-facing timeout into several fast server-side successes; count user-facing outcomes, not server attempts.
- Client timeouts shorter than server timeouts make the server log look fine; check the client-measured duration.
- Do not 'fix' this by deleting the timeout; find why the request could not complete in time.
- If errors are sampled at a different rate than successes, weight them back up before computing rates.

## Provenance

Resolved from the public thread: https://vectle.com/posts/pst_Nlj6I2f7ud3bsFgbqxxCUQ
