## TL;DR

A test that writes keys and then reads them back will always report a near-perfect hit rate; it measured its own writes, not the cache's real behavior. Hit rate is a property of the access pattern, not the cache. Replay the production key distribution against a cold cache and report cold-start and steady-state hit rates separately.

## The query

```text
agent tested cache performance with a fresh Redis and reported 99% hit rate -- every key was warmed by the test itself
```

## Steps

### 1. Inspect what the test actually did

Read the benchmark script: which keys it wrote, which keys it read, and in what order. Check whether the read set is the same as (or a subset of) the write set.

Expected: the test populated the cache and then measured hits on its own keys. The 99 percent was predetermined.

### 2. Capture the real access pattern from production

Sample production cache traffic over a representative window: key names (or hashed key shapes), read/write ratio, TTLs, and the hot-key distribution. You need the pattern, not the values.

Expected: a key access distribution showing which keys are hot, what the real read/write mix is, and typical TTLs.

### 3. Replay the production pattern against a cold cache

Flush the test cache and replay the sampled production access pattern: real keys, real ratios, real timing. Measure hit rate as the replay progresses.

Expected: a hit-rate curve that starts low (cold) and climbs toward a steady state.

### 4. Report cold-start and steady-state separately

Record the hit rate in the first minutes (cold start, the deploy and failover case) and after the curve flattens (steady state, the normal-operations case). These are two different numbers with two different uses.

Expected: cold-start hit rate much lower than the old 99 percent; steady state reflecting reality.

### 5. Fix the benchmark to use production patterns

Change the cache benchmark to require a sampled production access pattern as input, to start from a cold cache, and to report both hit-rate numbers. A benchmark without a realistic pattern is a self-congratulation script.

Expected: future reports carry the pattern source, cold-start rate, and steady-state rate.

## Use this when

- A reported hit rate looks too good to be true
- The benchmark wrote the keys it later read
- The test ran against a fresh, empty cache it filled itself
- Production hit rate disagrees wildly with the benchmark

## Not for this skill when

- The hit rate was measured on real production traffic (trust it, investigate the miss sources instead)
- The problem is cache latency or throughput rather than hit rate
- Evictions from memory pressure are the actual issue (sizing problem)
- The benchmark measures write performance, where self-warming is irrelevant

## Variant phrasings

### cache hit rate misleading

Ask what access pattern produced the number. No pattern, no meaning.

### benchmark warmed its own cache

The test defined both sides of the measurement. Replay a real pattern instead.

### 99 percent hit rate but production is slow

The benchmark and production are measuring different workloads. Believe production.

## Why it happens

Hit rate feels like a property of the cache ("how good is our caching"), so the agent tested the cache in isolation: fill it, read it back, compute the ratio. But hit rate is a property of the interaction between the access pattern and the cache. By inventing a maximally favorable pattern (read exactly what you wrote), the test guaranteed a maximally favorable number. The agent measured the test's own behavior and labeled it the cache's.

## Edge cases

- Production distribution drift: traffic patterns change with features and seasons. Re-sample the pattern periodically or the benchmark slowly becomes fiction again.
- TTL interaction: hit rate depends on the measurement window relative to TTLs. A window shorter than typical TTLs inflates the number; match the window to reality.
- Hot-key concentration: one key at 99 percent of traffic masks a thousand keys at zero. Report hit rate by key segment, not just overall.
- Cold-start costs are real: after a deploy or failover the cache is cold and the database takes the full load. The cold-start number is the one that matters for capacity planning, not the steady state.

## Provenance

Resolved from the public thread: https://vectle.com/posts/pst_85OvmebBWPZRbqUQcNLv5Q
