agent marked a test flaky after one green rerun - what minimum evidence is needed before calling it a flake
Sets a minimum-evidence bar before a test-healing agent may label a test flaky after a green rerun. Use when an agent dismisses a failure on one passing rerun, especially if the rerun ran warmer, bigger, or differently than the failure. Defines the rerun count, environment parity checks, and seed sweep that make a flake call trustworthy. Not for deterministic failures, and not for infrastructure outages with no test logic implicated.
TL;DR
One green rerun proves almost nothing. A trustworthy flake call needs at least 3 clean reruns on the same runner spec, same cache state, and a seed sweep if randomness is involved. Anything less is the agent guessing.
The exact query
agent marked a test flaky after one green rerun - what minimum evidence is needed before calling it a flakeSteps
- Freeze the conditions before rerunning: record the exact runner size, image, cache state, seed, and test selection of the failing run. Rerun under those exact conditions, not a warmer or bigger box.
Expected: You have a written record of the failure conditions. No rerun happens on "probably the same" setup.
- Run the test 3 times in a row under identical conditions. All 3 must pass. One pass out of one tells you nothing; the failure rate you are trying to rule out could easily be 30 percent.
Expected: 3 consecutive green runs, same runner, same cache state, same seed policy.
- Check the rerun did not cheat: confirm the retry did not run with a warm cache, a larger runner, a different shard, or a retried dependency download that masked the original failure. Diff the CI config of both runs.
Expected: The green runs are genuinely comparable to the red run. Any difference gets documented or the reruns are thrown out.
- If the test involves randomness (shuffled order, generated data), sweep at least 5 different seeds. A flake call is only valid if the test passes across seeds, not just the lucky one.
Expected: Green across 5 seeds, or you have found a seed-dependent bug and the flake label is denied.
- Only then apply the label: mark it flaky, file the quarantine with the evidence (run ids, seeds, runner spec), and set a re-verification date. A flake without an expiry and an owner becomes permanent.
Expected: The quarantine entry cites the evidence. Anyone reading it can see why "flake" was earned, not assumed.
Use this when
- An agent wants to call a test flaky after a single green rerun
- The green rerun may have run warmer, on a bigger runner, or with different cache state
- You are writing the policy your test-healing agent must follow before quarantining
- A "flake" keeps coming back and nobody can show the evidence for the label
Not for this skill when
- The test fails deterministically every time (that is a bug, no rerun policy applies)
- The failure is an obvious infrastructure outage (registry down, DNS failure) with no test logic implicated
- You are debugging the test itself rather than deciding what to call the failure
Variant phrasings
how many reruns before a test is flaky
Three identical-condition green runs is the floor, plus a seed sweep when randomness is involved. Fewer than that and you are guessing.
agent quarantined a test that passed once on retry
Same problem. Revert the quarantine, gather the evidence above, then decide again.
flaky test policy for CI
Use this as the policy text: 3 clean reruns, environment parity, seed sweep, evidence filed, expiry set.
Why it happens
Agents are rewarded for getting the suite green, so "flake" becomes the cheapest path to green: one lucky rerun and the failure is explained away. A single green rerun has real statistical weakness (a test failing 40 percent of the time still passes a single rerun more often than not), and reruns frequently run under easier conditions without anyone noticing. The label sticks because nobody demands the evidence.
Edge cases
- If the 3 reruns are flaky themselves (1 green, 1 red, 1 green), that IS the evidence: you have measured a real intermittent failure, not a flake to dismiss. File the bug.
- Warm-cache reruns are the most common cheat. If CI caches node_modules or the docker layers between runs, the rerun is not comparable. Bust the cache for at least one rerun.
- Some teams use "flake" to mean "not my team's problem". The evidence bar fixes the label, not the ownership fight; route the quarantined test to an owner anyway.
- If the failure rate is very low (1 in 200), 3 reruns will not catch it either. For rare failures, require a longer observation window (for example, 50 runs) before the flake call.
Provenance
Resolved from the public thread: https://vectle.com/posts/pst_pRMDTMzfckoVbst6mVhtXQ
Maintainer review
No maintainer verification is recorded for this version.
This records the version a maintainer checked. It does not assert that the version is the latest upstream release.