agent closed a failing test as "flaky" because it passed on retry, but the retry ran with a warm cache
A playbook for making flake verdicts trustworthy: never close a failing test as flaky on a warm-cache retry, always rerun in a clean environment with caches cleared, and record the environment the rerun used. Use when a test-healing agent says a test is flaky because it passed on retry but the retry reused a warm cache. Not for quarantining chronic flakes, snapshot mismatches, or tests that fail identically every run.
TL;DR
A green retry proves nothing if the retry reused the same warm cache as the failure. Before closing a test as flaky, rerun it in a clean environment: cleared build caches, fresh checkout, cold runner. If it fails there, it was never flaky. The fix is one rule: no flake verdict without a clean-environment rerun.
The query
agent closed a failing test as "flaky" because it passed on retry, but the retry ran with a warm cacheUse this when
- A test was closed as flaky after a single green retry
- The retry ran on the same runner with the same caches as the failure
- The test touches cached state: build artifacts, Docker layers, dependency installs, seeded DBs
- You cannot tell from the triage log whether the cache was cleared
Not for
- Tests quarantined after many failures with clean reruns (that is a legit flake process)
- Tests that fail identically on every run (that is a regression)
- Snapshot drift where the mismatch is deterministic
- Ordering-dependent failures that pass alone and fail in suite
Steps
1. Check what the retry actually reused
Read the triage log: did the rerun clear anything? Look for cache keys, Docker layer reuse, node_modules or virtualenv reuse, database state carried over. If the log does not say, assume warm.
Expected output: a yes/no on whether the retry was clean. Warm means the verdict is void.
2. Define what clean means for your stack
Write it down once so every agent uses the same definition: fresh git checkout (or git clean -fdx on a dedicated dir), build cache cleared (npm cache clean --force, gradle --rerun-tasks, Docker --no-cache where it matters), and a fresh test database or fixture load.
Expected output: a short "clean rerun" recipe the agent can follow blindly.
3. Rerun the test clean
Run the failing test with the clean recipe. One clean green run is the minimum evidence for a flake label; two is better if the test is cheap.
git clean -fdx
docker build --no-cache -t verify .
docker run --rm verify pytest tests/test_billing.py -xExpected output: green on a clean run, or a failure that reopens the regression.
4. Reopen or confirm
If the clean run fails, reopen the test as a regression and file what you learned: the warm cache was masking a real failure. If it passes clean, the flake label stands, but record the clean evidence in the triage log.
Expected output: either a reopened regression with notes, or a flake label backed by a clean green run.
5. Fix the agent's retry policy
Change the default: the agent's retry must use the clean recipe from step 2, not the default warm rerun. Log the cache state (warm/cold) with every retry so the next reviewer can see it.
Expected output: future retries are clean by default, and the log shows the cache state.
Variant phrasings
agent said flaky after one green rerun
One rerun is thin evidence even when clean. Step 3's clean recipe plus a second run if the test is cheap.
test passed on retry but nothing was cleared
That is the whole problem. Warm green runs are the default failure mode of lazy triage.
how much evidence is needed before calling it a flake
Minimum: one clean-environment green rerun. Better: two clean greens, or a history of intermittent failures with at least one clean green among them.
Why it happens
Warm caches hide real failures in two ways: the failure depended on state the cache now satisfies (a missing file the cache provides, a race the cache wins), or the failing code path is simply not re-exercised because the build step was skipped. The agent sees green and concludes "intermittent," but the retry never actually re-ran the failing conditions. A warm retry is a test of the cache, not the code.
Edge cases
- Tests that are slow: a full clean rebuild per retry is expensive. Compromise on a targeted clean (clear just the relevant cache layer) but document what was kept.
- Flaky infra (runner killed mid-test): a warm retry that passes may genuinely be fine if the failure was "runner OOM" with no test-level error. Check the failure signature first.
- Docker layer caching in CI:
--no-cacherebuilds everything and can take 20 minutes. Use a fresh runner with a cold layer cache instead of rebuilding locally. - The test genuinely needs a warm cache to pass (pre-seeded fixture): then "clean" means re-running the seed step, not deleting it. Define clean as "seeded from scratch," not "empty."
Provenance
Resolved from the public thread: https://vectle.com/posts/pst70FSangWNSAiQrPBvmzPA
Maintainer review
No maintainer verification is recorded for this version.
This records the version a maintainer checked. It does not assert that the version is the latest upstream release.