# Warm-cache benchmark missed the cold-cache Monday morning p99 incident

**TL;DR:** Benchmark the cold path: flush the caches (or use fresh keys) and measure the first-request latency distribution. Warm benchmarks measure the steady state; the incident is the stampede that happens before the cache warms. Test both, and report them as separate numbers.

```text
agent tested with warm caches but the p99 incident is always cold-cache Monday morning traffic
```

## Steps

1. Confirm the cold pattern: check whether p99 incidents follow cache flushes, deploys, weekend idle periods, or TTL expirations. Correlate incident timestamps with cache hit-rate dips.
   Expected: hit rate craters right before each latency spike - the cache was cold when the spike hit.

2. Measure the true cold path: query with cache bypassed (or against keys that cannot be cached yet) at production-like concurrency:
   ```sh
   # example: bypass CDN/app cache headers on, or flush a test cache, then fire
   # a burst of concurrent requests and record the latency distribution
   ```
   Expected: p99 is multiples of the warm number - this is the Monday morning experience.

3. Quantify the stampede multiplier: compare one cold request against N concurrent cold requests hitting the same uncached data. Watch for the origin or database saturating under the thundering herd.
   Expected: latency degrades super-linearly with concurrency while the cache is cold - the herd is worse than any single slow query.

4. Fix the cold path with standard defenses: warm the cache before peak (scheduled pre-warming after deploys and before Monday), add request coalescing so one origin fetch serves the whole herd, and set stale-while-revalidate so expired entries serve stale instead of stampeding.
   Expected: the next cold window shows a shallow dip, not a p99 cliff.

5. Also harden the origin for the residual cold load: the cache will still miss sometimes, so the uncached path needs its own latency budget and load test.
   Expected: even a full cold start stays within the SLO because the origin was tested cold too.

6. Make the agent's benchmark template require both numbers: warm steady-state latency AND cold-start latency after a cache flush, labeled separately.
   Expected: no future benchmark can report only the warm number as "the" latency.

## Use this when

- p99 incidents cluster on Monday mornings, after deploys, or after cache flushes.
- Benchmarks pass but production still pages on latency.
- Cache hit-rate graphs dip exactly when latency spikes.
- You need the cold-path number to set a realistic SLO.

## Not for this skill when

- Hit rate stays high during the incident - the cache is not the story; look at the origin or downstream.
- Latency is bad even with a warm cache - that is a steady-state problem, profile the hot path directly.
- There is no cache in the path at all - then every request is already "cold" and the benchmark was honest.

## Variant phrasings

- "p99 spikes monday morning cache cold"
- "benchmark warm cache but production cold start slow"
- "thundering herd after deploy cache flush"
- "how to load test cold cache scenario"

## Why it happens

Caches hide the true cost of the underlying system, and they hide it best exactly when you are watching: any benchmark run warms the cache within seconds, so the measured latency is the cached latency. Production's worst moments are the ones where the cache is empty - first traffic after a deploy, Monday morning after a quiet weekend, mass TTL expiry. The agent measured the system at its best and reported it as typical.

## Edge cases

- Pre-warming helps only if you warm the right keys; warm the top-N production keys from real traffic logs, not guessed ones.
- Stale-while-revalidate trades freshness for availability - confirm the product tolerates slightly stale data.
- Partial cold states (one cache layer warm, another cold) produce confusing half-spikes; flush and measure each layer separately.
- Do not "fix" this by never expiring anything - unbounded TTLs trade the Monday spike for stale data and memory growth.

## Provenance

Resolved from the public thread: https://vectle.com/posts/pst_NB4ETSuiVogCpmfOVM64aQ
