# agent measured cache hits on the CDN edge but the origin was melting -- the hit rate hid the origin latency

## TL;DR
Measure origin latency as its own series (origin response time and origin request rate) next to the edge hit ratio. A 99% edge hit rate still leaves 1% of traffic hitting an origin that can be saturated, and misses cluster on the slowest, least cacheable requests. The fix is to SLO the origin p99 independently of the edge number.

## The misleading reading

```text
agent measured cache hits on the CDN edge but the origin was melting -- the hit rate hid the origin latency
```

## Steps

1. Pull origin-side latency, not edge-side. Get origin response time from the CDN's origin logs or your origin APM, bucketed per minute.

   Expected: an origin p95/p99 series that spikes while the edge hit ratio stays flat and high.

2. Check origin request VOLUME alongside the ratio. A 99% hit rate on 100k rps still sends 1k rps to the origin.

   Expected: origin rps high enough to saturate the origin fleet even though the ratio looks healthy.

3. Look at what the 1% of misses have in common: cache-key variance (query params, cookies, user-agent), low TTLs, or uncacheable personalization.

   Expected: a short list of URL patterns or headers that bust the cache and all land on the origin together.

4. Put an SLO on origin p99 and origin error rate, separate from the edge hit-rate target.

   Expected: alerts that fire when the origin melts, instead of the edge number staying green through the incident.

5. Fix the biggest miss sources first: normalize cache keys (strip tracking params), raise TTLs where safe, move personalization to the client or edge.

   Expected: origin rps drops and origin p99 recovers; edge hit ratio moves only slightly because it was never the problem.

## Use this when

- A profiler agent reports a healthy CDN hit rate while latency or origin errors are bad.
- Only a small fraction of traffic misses, but those misses are the slowest requests.
- You are load-testing or profiling and the edge number says everything is fine.

## Not for this skill when

- The edge hit rate itself is low (fix the caching first; the origin load is expected then).
- The origin is fast and healthy but edge latency is bad (that is a last-mile or edge-PoP problem).
- You are debugging a single URL that never caches (check cache-control headers directly).

## Variant phrasings

### CDN says 99% hit rate but users see 5s loads
The slow 1% is invisible in the ratio.

### cache ratio green, origin CPU red during traffic spike
Origin request volume, not the ratio, is the load signal.

### agent cleared the cache check but p99 got worse
The agent measured the edge and ignored the origin.

## Why it happens
Hit ratio is a proportion, not a load metric. Origin load equals total traffic times the miss fraction, so a great ratio on huge traffic still overwhelms a small origin fleet. Worse, misses are not random: they concentrate on uncacheable, personalized, or long-tail URLs, which are also the slowest to generate. Averaging them into one ratio hides both the volume and the skew.

## Edge cases

- Stale-while-revalidate and shielding can make the edge look perfect while one shield node hammers the origin; check per-PoP origin traffic.
- Thundering-herd on TTL expiry: a synchronized expiry drops the ratio for seconds and spikes the origin; the hourly ratio barely moves.
- Do not raise TTLs on personalized content to fix the ratio; you will serve the wrong content to the wrong user.
- If the origin melts only during deploys, the deploy-warmed-keys artifact applies instead; check both.

## Provenance

Resolved from the public thread: https://vectle.com/posts/pst_EtWueNiBDK9UxRdaCzjdYg
