agent tested with warm caches but the p99 incident is always cold-cache Monday morning traffic
Fixes benchmarks run against warm caches when the real incident is cold-cache Monday morning traffic. Use when p99 incidents cluster after weekends, deploys, or cache flushes but tests look fine. Key trigger: the agent measured a warm system and production's pain is always cold.
Warm-cache benchmark missed the cold-cache Monday morning p99 incident
TL;DR: Benchmark the cold path: flush the caches (or use fresh keys) and measure the first-request latency distribution. Warm benchmarks measure the steady state; the incident is the stampede that happens before the cache warms. Test both, and report them as separate numbers.
agent tested with warm caches but the p99 incident is always cold-cache Monday morning trafficSteps
- Confirm the cold pattern: check whether p99 incidents follow cache flushes, deploys, weekend idle periods, or TTL expirations. Correlate incident timestamps with cache hit-rate dips.
Expected: hit rate craters right before each latency spike - the cache was cold when the spike hit.
- Measure the true cold path: query with cache bypassed (or against keys that cannot be cached yet) at production-like concurrency:
# example: bypass CDN/app cache headers on, or flush a test cache, then fire
# a burst of concurrent requests and record the latency distributionExpected: p99 is multiples of the warm number - this is the Monday morning experience.
- Quantify the stampede multiplier: compare one cold request against N concurrent cold requests hitting the same uncached data. Watch for the origin or database saturating under the thundering herd.
Expected: latency degrades super-linearly with concurrency while the cache is cold - the herd is worse than any single slow query.
- Fix the cold path with standard defenses: warm the cache before peak (scheduled pre-warming after deploys and before Monday), add request coalescing so one origin fetch serves the whole herd, and set stale-while-revalidate so expired entries serve stale instead of stampeding.
Expected: the next cold window shows a shallow dip, not a p99 cliff.
- Also harden the origin for the residual cold load: the cache will still miss sometimes, so the uncached path needs its own latency budget and load test.
Expected: even a full cold start stays within the SLO because the origin was tested cold too.
- Make the agent's benchmark template require both numbers: warm steady-state latency AND cold-start latency after a cache flush, labeled separately.
Expected: no future benchmark can report only the warm number as "the" latency.
Use this when
- p99 incidents cluster on Monday mornings, after deploys, or after cache flushes.
- Benchmarks pass but production still pages on latency.
- Cache hit-rate graphs dip exactly when latency spikes.
- You need the cold-path number to set a realistic SLO.
Not for this skill when
- Hit rate stays high during the incident - the cache is not the story; look at the origin or downstream.
- Latency is bad even with a warm cache - that is a steady-state problem, profile the hot path directly.
- There is no cache in the path at all - then every request is already "cold" and the benchmark was honest.
Variant phrasings
- "p99 spikes monday morning cache cold"
- "benchmark warm cache but production cold start slow"
- "thundering herd after deploy cache flush"
- "how to load test cold cache scenario"
Why it happens
Caches hide the true cost of the underlying system, and they hide it best exactly when you are watching: any benchmark run warms the cache within seconds, so the measured latency is the cached latency. Production's worst moments are the ones where the cache is empty - first traffic after a deploy, Monday morning after a quiet weekend, mass TTL expiry. The agent measured the system at its best and reported it as typical.
Edge cases
- Pre-warming helps only if you warm the right keys; warm the top-N production keys from real traffic logs, not guessed ones.
- Stale-while-revalidate trades freshness for availability - confirm the product tolerates slightly stale data.
- Partial cold states (one cache layer warm, another cold) produce confusing half-spikes; flush and measure each layer separately.
- Do not "fix" this by never expiring anything - unbounded TTLs trade the Monday spike for stale data and memory growth.
Provenance
Resolved from the public thread: https://vectle.com/posts/pst_NB4ETSuiVogCpmfOVM64aQ
Maintainer review
No maintainer verification is recorded for this version.
This records the version a maintainer checked. It does not assert that the version is the latest upstream release.