how to run a test 100 times to prove it is not flaky
Shows how to repeat a single test 100 times with go test -count, pytest-repeat --count, or a shell loop for rspec, and how to read the result as evidence rather than proof. Use after fixing a flake to build confidence, before making a test merge-blocking, or to measure a failure rate. Not as proof of correctness and not for slow or time-dependent tests without controls.
TL;DR
Repeat the test with your runner's repeat flag (go test -count=100, pytest --count=100 via pytest-repeat, or a shell loop for rspec) and require 100 out of 100 passes. One hundred green runs is evidence, not proof; pair it with -race or stress flags if the flake is concurrency-shaped.
Problem
You fixed a suspected flake (or suspect one) and want repeated-run evidence that the test is now stable.
Steps
- Pick the repetition method for your runner:
go test ./cart -run '^TestCheckout$' -count=100
pytest tests/test_checkout.py::test_discount --count=100 -q
for i in $(seq 1 100); do rspec spec/checkout_spec.rb || break; done (--count needs the pytest-repeat plugin: pip install pytest-repeat.) Expected: the command starts executing the test repeatedly.
- Watch for any single failure; stop on the first red and investigate that run's seed and output instead of continuing.
Expected: either 100 out of 100 green, or one failing run with captured output.
- If the flake was timing or concurrency related, add the stress flags too:
go test ./cart -run '^TestCheckout$' -count=100 -race
pytest tests/test_checkout.py::test_discount --count=100 -n 4Expected: still 100 out of 100 under the stressor that used to trigger the flake.
- Record the result with the exact command and the commit hash so the claim is checkable later.
Expected: anyone can rerun the same command on the same commit and see the same outcome.
When to use
- After fixing a flaky test, to build confidence before closing the issue.
- Before marking a test as reliable enough for merge-blocking CI.
- To characterize a suspected flake's failure rate.
When not to use
- As proof of correctness: 100 passes cannot prove absence of a race, only that it did not trigger.
- Slow suites: 100x a 30-second test is 50 minutes; sample fewer runs or parallelize.
- Tests that depend on real time or external services: repetition without controlling those variables proves little.
Tool compatibility
- Go 1.22 through 1.24:
-countbuilt in;-racebuilt in. - pytest 8.x:
--countvia the pytest-repeat plugin. - RSpec 3.x: no built-in repeat; use the shell loop shown above.
Variant phrasings
repeat a pytest test N times
Install pytest-repeat and pass --count=N; combine with -x to stop at the first failure.
stress test a single go test
go test -run '^TestName$' -count=100 -race is the standard one-liner.
Why it happens
Intermittent failures have a per-run probability; if the true failure rate were even 5 percent, the chance of 100 clean runs is under 1 percent. Repetition turns "it passed once" into a quantified claim.
Edge cases
- Order-dependent flakes: repeating one test in isolation misses pollution from other tests; also run the full file or suite a few times.
- Caching:
go testcaches results;-count=100disables the cache automatically, but plain reruns need-count=1. - The 101st run fails in CI: treat any single failure as signal, not noise; reopen the investigation.
Provenance
Resolved from the public thread: https://vectle.com/posts/pst_u7UgfIMXvnXxF1AsZFk4QQ
Maintainer review
No maintainer verification is recorded for this version.
This records the version a maintainer checked. It does not assert that the version is the latest upstream release.