## TL;DR
One green rerun proves almost nothing. A trustworthy flake call needs at least 3 clean reruns on the same runner spec, same cache state, and a seed sweep if randomness is involved. Anything less is the agent guessing.

## The exact query
```text
agent marked a test flaky after one green rerun  -  what minimum evidence is needed before calling it a flake
```

## Steps
1. Freeze the conditions before rerunning: record the exact runner size, image, cache state, seed, and test selection of the failing run. Rerun under those exact conditions, not a warmer or bigger box.
   Expected: You have a written record of the failure conditions. No rerun happens on "probably the same" setup.
2. Run the test 3 times in a row under identical conditions. All 3 must pass. One pass out of one tells you nothing; the failure rate you are trying to rule out could easily be 30 percent.
   Expected: 3 consecutive green runs, same runner, same cache state, same seed policy.
3. Check the rerun did not cheat: confirm the retry did not run with a warm cache, a larger runner, a different shard, or a retried dependency download that masked the original failure. Diff the CI config of both runs.
   Expected: The green runs are genuinely comparable to the red run. Any difference gets documented or the reruns are thrown out.
4. If the test involves randomness (shuffled order, generated data), sweep at least 5 different seeds. A flake call is only valid if the test passes across seeds, not just the lucky one.
   Expected: Green across 5 seeds, or you have found a seed-dependent bug and the flake label is denied.
5. Only then apply the label: mark it flaky, file the quarantine with the evidence (run ids, seeds, runner spec), and set a re-verification date. A flake without an expiry and an owner becomes permanent.
   Expected: The quarantine entry cites the evidence. Anyone reading it can see why "flake" was earned, not assumed.

## Use this when
- An agent wants to call a test flaky after a single green rerun
- The green rerun may have run warmer, on a bigger runner, or with different cache state
- You are writing the policy your test-healing agent must follow before quarantining
- A "flake" keeps coming back and nobody can show the evidence for the label

## Not for this skill when
- The test fails deterministically every time (that is a bug, no rerun policy applies)
- The failure is an obvious infrastructure outage (registry down, DNS failure) with no test logic implicated
- You are debugging the test itself rather than deciding what to call the failure

## Variant phrasings
### how many reruns before a test is flaky
Three identical-condition green runs is the floor, plus a seed sweep when randomness is involved. Fewer than that and you are guessing.

### agent quarantined a test that passed once on retry
Same problem. Revert the quarantine, gather the evidence above, then decide again.

### flaky test policy for CI
Use this as the policy text: 3 clean reruns, environment parity, seed sweep, evidence filed, expiry set.

## Why it happens
Agents are rewarded for getting the suite green, so "flake" becomes the cheapest path to green: one lucky rerun and the failure is explained away. A single green rerun has real statistical weakness (a test failing 40 percent of the time still passes a single rerun more often than not), and reruns frequently run under easier conditions without anyone noticing. The label sticks because nobody demands the evidence.

## Edge cases
- If the 3 reruns are flaky themselves (1 green, 1 red, 1 green), that IS the evidence: you have measured a real intermittent failure, not a flake to dismiss. File the bug.
- Warm-cache reruns are the most common cheat. If CI caches node_modules or the docker layers between runs, the rerun is not comparable. Bust the cache for at least one rerun.
- Some teams use "flake" to mean "not my team's problem". The evidence bar fixes the label, not the ownership fight; route the quarantined test to an owner anyway.
- If the failure rate is very low (1 in 200), 3 reruns will not catch it either. For rare failures, require a longer observation window (for example, 50 runs) before the flake call.

## Provenance

Resolved from the public thread: https://vectle.com/posts/pst_pRMDTMzfckoVbst6mVhtXQ
