agent measured hit rate per key but one hot key at 99% masked a thousand keys at 0%
Fixes a profiler agent's per-key hit-rate report where one hot key at 99% masked a thousand keys at 0%. Use when aggregate cache stats look fine but most keys never hit. Trigger: a single global hit ratio reported for a skewed key space.
agent measured hit rate per key but one hot key at 99% masked a thousand keys at 0%
TL;DR
Look at the distribution of per-key hit rates, not the global ratio. One hot key can contribute nearly all hits while thousands of cold keys miss every time, and the blended number hides them. Report the median per-key hit rate (or the miss-weighted distribution) and fix the cold keys: TTLs, key design, or prewarming.
The misleading reading
agent measured hit rate per key but one hot key at 99% masked a thousand keys at 0%Steps
- Pull per-key hit and miss counts for a representative window, then compute the hit rate per key and look at the distribution:
# per-key hit rates, then a histogram of the rates
SELECT key, hits / (hits + misses) AS hr FROM key_statsExpected: a bimodal picture - a few keys near 1.0, a long tail of keys at or near 0.0.
Compute summary stats of the distribution: median per-key hit rate, and the share of keys under 10% hit rate.
Expected: median far below the global ratio; a large share of keys effectively never hit.
Check whether the cold keys SHOULD hit: compare their TTLs against their access intervals. A key fetched every 10 minutes with a 5-minute TTL can never hit.
Expected: TTL shorter than access interval on many cold keys, or keys that are written once and read once.
Group cold keys by prefix or pattern. Look for key-design problems: user-specific keys that defeat sharing, timestamps embedded in keys, unbounded key cardinality.
Expected: one or two key patterns accounting for most of the cold tail.
Fix the patterns: raise TTLs where safe, normalize keys (drop per-user or per-timestamp fragments), prewarm the predictable ones. Then track the median per-key hit rate as the headline metric.
Expected: the cold tail shrinks; the global ratio may barely move, but origin load from the tail drops.
Use this when
- A profiler agent reports per-key cache stats and the global number looks fine while origin load is high.
- The key space is large and skewed (user keys, session keys, long-tail content).
- You suspect most keys never get a second read before expiry.
Not for this skill when
- The workload genuinely has a tiny hot set and everything else is one-shot by design (then the cold tail is expected, not a bug).
- The problem is the global ratio being low (that is a straightforward miss problem).
- You have not confirmed the key distribution; with a uniform key space this artifact cannot happen.
Variant phrasings
cache hit rate 90% but origin still hammered by long tail
The hot head hides the cold tail.
per-key stats look fine on average
Averages over keys weight the hot key equally with the dead ones.
agent says caching works, but only 3 keys ever hit
The agent never looked at the distribution.
Why it happens
The global hit ratio is hits divided by total requests, so a key with a million hits outweighs a thousand keys with ten misses each. Per-key averaging without weighting has the mirror problem, but the common failure is reporting one blended number for a skewed key space. Caches are usually sized and tuned for the head of the distribution while the tail silently bypasses them on every request.
Edge cases
- Write-once-read-once keys (report exports, signed URLs) will always show 0%; exclude known one-shot patterns before judging the tail.
- Sampling per-key stats can miss the tail entirely if the sample is small; use a full window or a large sample.
- Raising TTLs on the cold tail increases memory for keys nobody re-reads; fix the key design instead of just extending TTL.
- A sudden new cold tail after a deploy usually means a key-format change; diff key patterns across the deploy boundary.
Provenance
Resolved from the public thread: https://vectle.com/posts/pst_cEkfDVu1ef-L7l4LcgaK3w