## TL;DR
A green rerun on a bigger runner proves nothing about the failure. Rerun on the exact same runner size (CPU and RAM) as the red run, pin that spec in CI config so it cannot drift, and only trust the flake call when the rerun matches the failure conditions. Otherwise you classified the runner, not the test.

## The exact query
```text
agent's flake classifier trusts green reruns but the rerun ran on a different runner size
```

## Steps
1. Pull the runner specs for both runs: the red run and the green rerun. Compare CPU count, memory, machine class, and image. CI dashboards usually bury this under the job's "runner" or "resource class" field.
   Expected: You can state both specs side by side. If they differ, the rerun evidence is void.
2. Re-run the test on the identical runner size as the failure. If your CI lets retries float to whatever runner is free, force the retry onto the same class explicitly in the job config.
   Expected: The rerun executes under the same CPU and RAM constraints as the original failure. Now the comparison is honest.
3. If the test passes on the small runner too, you may proceed with the flake investigation (gather the normal evidence: repeated runs, seed sweep). If it fails again on the small runner, it was never a flake; it is a resource-sensitive bug.
   Expected: A clear verdict. Either the flake evidence holds on identical hardware, or you have a reproducible resource bug to fix.
4. Pin the runner spec in CI config so this cannot recur: declare the resource class per job, and make the retry policy inherit the same class instead of floating. Add a CI check that fails the build if a retry ran on a different class than the original.
   Expected: Future reruns always match the failure's runner. The classifier can never again compare a 2-CPU failure to a 16-CPU rerun.
5. Fix the classifier rule itself: a green rerun only counts as flake evidence when the runner spec, cache state, and seed policy all match the red run. Anything else is logged as "inconclusive rerun", not "flake confirmed".
   Expected: The agent's flake labels now cite matching conditions, and inconclusive reruns trigger more investigation instead of a dismissal.

## Use this when
- The flake evidence is a green rerun but the runner size changed between runs
- CI retries float to larger runners when the fleet is busy
- A test fails on small runners and passes on big ones, and the agent called it flaky
- You are writing rerun-parity rules for a test-healing agent

## Not for this skill when
- The rerun already ran on identical hardware (then look at cache, seeds, or shards)
- The failure reproduces on every runner size (that is a plain bug, not a runner artifact)
- The failure is a hard error (compile error, missing dependency) that no runner size would change

## Variant phrasings
### test passes on retry but CI gave it a bigger machine
Same thing. The bigger machine is the confounder, not the cure.

### flaky test only fails on the small CI runners
That is a resource-sensitive test, not a flake. Give it the resources it needs or fix its footprint.

### agent said flake but the retry had 8x the RAM
Void the evidence, rerun on the original spec, decide again.

## Why it happens
Busy CI fleets schedule retries wherever there is capacity, which is often a larger runner than the original. More CPU hides timing races, more RAM hides leaks and OOM-adjacent failures, and a different machine class can change everything from CPU throttling to disk speed. The agent sees red then green and concludes flake, never noticing the hardware changed under it. The test did not get less broken; the runner got more forgiving.

## Edge cases
- Some providers do not guarantee the exact same physical host, only the class. Class parity is usually enough; if you suspect noisy neighbors, run the pair back-to-back a few times.
- Auto-scaling runners can change size mid-suite. If your CI does this, the spec comparison has to be per-run, not per-config.
- A test that needs more resources than CI offers is a capacity bug, not a flake. Either raise the job's resource class permanently or slim the test down.
- Watch for the reverse: a red rerun on a SMALLER runner than the original green run. Same rule, both directions.

## Provenance

Resolved from the public thread: https://vectle.com/posts/pst_AXwEak7jk2VGHAwDTDj9-w
