the agent said "flake" because the error changed every run, but it was actually an uninitialized random seed
Fixes test-healing agents that label a changing-error failure a flake when the real cause is an uninitialized random seed. Use when a test fails with a different error message on every run, especially under pytest-randomly or hypothesis, and reruns keep producing new failure modes. Pin the seed to reproduce the stable failure, then fix the real ordering or randomness bug underneath. Not for infrastructure flakes that fail with the same error, and not for tests with no randomness involved.
TL;DR
A different error on every run is not a flake, it is nondeterminism the test never pinned down. Set PYTHONHASHSEED and the runner's seed to a fixed value, reproduce one stable failure, and fix that. The changing errors were just one bug wearing different masks.
The exact query
the agent said "flake" because the error changed every run, but it was actually an uninitialized random seedSteps
- Confirm randomness is in play: check whether the suite uses pytest-randomly, hypothesis, factory_boy, or faker. Run the failing test 5 times with a fixed seed (for pytest-randomly, pass -p no:randomly with a fixed seed value via --randomly-seed=42).
Expected: With a fixed seed, the failure is the same error every time. If it is, you have a deterministic bug, not a flake.
- Lock down every seed source: set the PYTHONHASHSEED environment variable to 0, seed the language RNG (random.seed in python, Math.random is not seedable so look for a seeded PRNG in JS), and pass --randomly-seed in CI. Record the seed that reproduces the failure.
Expected: The same seed reproduces the same failure on demand, on any machine.
- Now debug the single stable failure like any normal bug. The usual culprits are iteration over an unordered set or dict, a test that depends on factory-generated data values, or shared state ordered differently each run.
Expected: You find one root cause that explains all the different errors you saw before (for example, a dict iterated in random order feeding different rows to the assertion each run).
- Fix the real bug (sort the iteration, make the data deterministic, remove the order dependence), then re-run with several different seeds to prove the fix holds.
Expected: The test passes under seed 1, 42, 999, and whatever CI picks. The "flake" label is deleted.
- Make the seed policy permanent: CI always logs the seed it used, and the pipeline stores failing seeds so any rerun reproduces exactly.
Expected: The next time someone says "error changed every run", the seed log answers in one minute.
Use this when
- The error message is different on nearly every failing run
- pytest-randomly, hypothesis, faker, or factory-generated data is in the mix
- Reruns "fix" it, but the next failure looks nothing like the last one
- PYTHONHASHSEED or the runner seed was never set in CI
Not for this skill when
- The test fails with the same error every time (that is a plain bug, seed work buys you nothing)
- The failure is clearly infrastructure (runner OOM, network blip, docker pull failing)
- There is no randomness in the test path at all (seeds will not change anything)
Variant phrasings
test passes or fails randomly with a different stack trace each time
Same root cause family. Fix the seed first, then read the stable trace.
hypothesis found a falsifying example that vanishes on rerun
Capture the failing example from the hypothesis database or the printed seed, replay it with --hypothesis-seed, then shrink and fix. Do not mark it flaky.
pytest-randomly failures that disappear when run in isolation
Isolation changes the seed stream. Reproduce with the exact seed CI logged, not by running the test alone.
Why it happens
Python hash randomization (PYTHONHASHSEED) and test-order shufflers deliberately randomize iteration order and execution order to surface order-dependent bugs. When a test has an order dependence or consumes random data without seeding, each run explores a different path and trips over a different symptom. An agent that only looks at the error text sees N different errors and concludes "flaky". The errors were never the signal; the unseeded randomness was.
Edge cases
- If fixing the seed still gives different errors, there are two bugs or a real concurrency issue (threads, async races). Seed pinning rules out the deterministic half first.
- Hypothesis stores failing examples in its example database; deleting that directory throws away your reproduction. Back it up before touching it.
- Some JS setups cannot seed Math.random at all; if the randomness lives in the app code, the fix is a seedable PRNG, not a test flag.
- Watch for seeds that only reproduce on one platform (dict ordering differs across python versions). Pin the python version in CI too.
Provenance
Resolved from the public thread: https://vectle.com/posts/pst_dn9CdCIzcSZdsCmmiUC3Tw
Maintainer review
No maintainer verification is recorded for this version.
This records the version a maintainer checked. It does not assert that the version is the latest upstream release.