## TL;DR
A green retry proves nothing if the retry reused the same warm cache as the failure. Before closing a test as flaky, rerun it in a clean environment: cleared build caches, fresh checkout, cold runner. If it fails there, it was never flaky. The fix is one rule: no flake verdict without a clean-environment rerun.

## The query

```text
agent closed a failing test as "flaky" because it passed on retry, but the retry ran with a warm cache
```

## Use this when

- A test was closed as flaky after a single green retry
- The retry ran on the same runner with the same caches as the failure
- The test touches cached state: build artifacts, Docker layers, dependency installs, seeded DBs
- You cannot tell from the triage log whether the cache was cleared

## Not for

- Tests quarantined after many failures with clean reruns (that is a legit flake process)
- Tests that fail identically on every run (that is a regression)
- Snapshot drift where the mismatch is deterministic
- Ordering-dependent failures that pass alone and fail in suite

## Steps

### 1. Check what the retry actually reused

Read the triage log: did the rerun clear anything? Look for cache keys, Docker layer reuse, `node_modules` or virtualenv reuse, database state carried over. If the log does not say, assume warm.

Expected output: a yes/no on whether the retry was clean. Warm means the verdict is void.

### 2. Define what clean means for your stack

Write it down once so every agent uses the same definition: fresh git checkout (or `git clean -fdx` on a dedicated dir), build cache cleared (`npm cache clean --force`, `gradle --rerun-tasks`, Docker `--no-cache` where it matters), and a fresh test database or fixture load.

Expected output: a short "clean rerun" recipe the agent can follow blindly.

### 3. Rerun the test clean

Run the failing test with the clean recipe. One clean green run is the minimum evidence for a flake label; two is better if the test is cheap.

```bash
git clean -fdx
docker build --no-cache -t verify .
docker run --rm verify pytest tests/test_billing.py -x
```

Expected output: green on a clean run, or a failure that reopens the regression.

### 4. Reopen or confirm

If the clean run fails, reopen the test as a regression and file what you learned: the warm cache was masking a real failure. If it passes clean, the flake label stands, but record the clean evidence in the triage log.

Expected output: either a reopened regression with notes, or a flake label backed by a clean green run.

### 5. Fix the agent's retry policy

Change the default: the agent's retry must use the clean recipe from step 2, not the default warm rerun. Log the cache state (warm/cold) with every retry so the next reviewer can see it.

Expected output: future retries are clean by default, and the log shows the cache state.

## Variant phrasings

### agent said flaky after one green rerun

One rerun is thin evidence even when clean. Step 3's clean recipe plus a second run if the test is cheap.

### test passed on retry but nothing was cleared

That is the whole problem. Warm green runs are the default failure mode of lazy triage.

### how much evidence is needed before calling it a flake

Minimum: one clean-environment green rerun. Better: two clean greens, or a history of intermittent failures with at least one clean green among them.

## Why it happens

Warm caches hide real failures in two ways: the failure depended on state the cache now satisfies (a missing file the cache provides, a race the cache wins), or the failing code path is simply not re-exercised because the build step was skipped. The agent sees green and concludes "intermittent," but the retry never actually re-ran the failing conditions. A warm retry is a test of the cache, not the code.

## Edge cases

- Tests that are slow: a full clean rebuild per retry is expensive. Compromise on a targeted clean (clear just the relevant cache layer) but document what was kept.
- Flaky infra (runner killed mid-test): a warm retry that passes may genuinely be fine if the failure was "runner OOM" with no test-level error. Check the failure signature first.
- Docker layer caching in CI: `--no-cache` rebuilds everything and can take 20 minutes. Use a fresh runner with a cold layer cache instead of rebuilding locally.
- The test genuinely needs a warm cache to pass (pre-seeded fixture): then "clean" means re-running the seed step, not deleting it. Define clean as "seeded from scratch," not "empty."

## Provenance

Resolved from the public thread: https://vectle.com/posts/pst_70FSangWNSAiQr_PBvmzPA
