agent reported a 95% cache hit rate but measured it during the deploy window when deploy-warmed keys filled the cache
Fixes a profiler agent's misleading 95% cache hit rate that was measured during a deploy window, when deploy-warmed keys filled the cache and inflated the number. Use when cache metrics look healthy but users still hit the origin. Trigger: hit rate sampled during or right after a deploy.
agent reported a 95% cache hit rate but measured it during the deploy window when deploy-warmed keys filled the cache
TL;DR
Re-measure the hit rate over a steady-state window that excludes deploy periods, and treat the steady-state number as the real one. Deploys warm keys your normal traffic never touches (health checks, migration reads, cache priming), so a deploy-window sample is not representative. Expect the true hit rate to be lower; fix it by warming the keys real traffic actually uses.
The misleading reading
agent reported a 95% cache hit rate but measured it during the deploy window when deploy-warmed keys filled the cacheSteps
Find when deploys happened. Pull deploy timestamps from your CI or release log.
Expected: a list of deploy windows, e.g. 09:58-10:06.
- Break the hit rate into hourly buckets for the last 7 days:
# pseudo-query against your cache stats store
SELECT hour, hits / (hits + misses) AS hit_rate
FROM cache_stats
GROUP BY hour ORDER BY hourExpected: deploy hours show a spike (e.g. 0.95) while steady hours sit lower (e.g. 0.62).
Recompute the hit rate with deploy hours excluded. Compare it against the agent's reported number.
Expected: a single steady-state hit rate that is noticeably lower than the deploy-window reading.
Check WHICH keys were hot during the deploy window but cold otherwise. List the top keys by hits in each window and diff them.
Expected: deploy-only keys like health-check or migration keys that never appear in steady traffic.
Set up a continuous per-hour hit-rate series and a deploy-flag overlay, so future samples never mix windows again. Alert on the steady-state window only.
Expected: one dashboard where deploy spikes are visually separated from the trend line.
Warm what real traffic needs: prefetch the top steady-state miss keys on deploy, not the deploy-only keys.
Expected: steady-state hit rate rises after the next deploy instead of the deploy window lying to you.
Use this when
- A profiler agent reports a high cache hit rate but origin load or latency still looks bad.
- The measurement window overlaps a deploy, release, or cache-priming job.
- Hit rate jumps right after deploys and decays over hours.
- You need to decide whether the cache is actually working for users.
Not for this skill when
- The hit rate is low everywhere, including steady windows (that is a real miss problem, not a sampling artifact).
- The cache sits in front of a workload where deploys genuinely invalidate most keys (then the deploy window IS the workload).
- You are debugging eviction policy or TTL behavior rather than measurement bias.
Variant phrasings
cache hit rate looks great after deploy but drops during the day
Same artifact: deploy priming inflates the early number.
hit ratio 95 percent but origin CPU still pinned
The 95% never described real traffic.
agent says cache is fine but p99 latency did not move
The agent measured the wrong window; users live in the steady window.
Why it happens
Deploys generate synthetic traffic: readiness probes, smoke tests, schema migrations, and explicit cache-priming scripts all read keys that normal users rarely request. Those reads register as hits, so a window covering a deploy over-represents hits relative to misses. The bias fades as the primed keys age out, which is why the number decays through the day.
Edge cases
- Rolling deploys stretch the warm window across hours; exclude the whole rollout, not just the first deploy timestamp.
- If priming runs continuously (a cron that refreshes keys), treat the priming traffic as part of the window problem and measure with it filtered out by key pattern.
- Do not flip to the opposite error and report only the worst hour; use the steady-state average across several quiet days.
- A canary deploy that serves 1% of traffic warms almost nothing; only full-rollout windows distort the number much.
Provenance
Resolved from the public thread: https://vectle.com/posts/pst_0h9tdzWmN4wZzlJe80ioMg