my agent quarantined a test that failed once in 200 runs - it turned out to be a real race condition, not noise
Teaches an agent how to tell a true one-in-200 race condition from random noise before quarantining, and how to recover when a real bug got quarantined. Use when a rarely failing test was auto-quarantined and a later investigation found a real concurrency bug. Not for bulk-quarantine cleanup, deterministic failures mislabeled as flaky, or tests that fail on every run.
TL;DR
A test that fails once in 200 runs can still be a real bug. Race conditions are rare by nature, so a low failure rate is exactly what a genuine race looks like. Before quarantining, the agent must stress the test (repeated runs under tightened timing) to distinguish a real race from infrastructure noise, and the quarantine decision must record that evidence. If a quarantined test hid a race, un-quarantine it, reproduce the race deterministically, and fix the synchronization, not the test.
The query
my agent quarantined a test that failed once in 200 runs - it turned out to be a real race condition, not noiseUse this when
- A rarely failing test was auto-quarantined and a later investigation found a real concurrency bug.
- The agent's quarantine policy keys only on failure rate.
- You need a recovery plan for a quarantine that buried a real bug.
Not for
- Bulk-quarantine cleanup (dozens of stale quarantines at once).
- Deterministic failures mislabeled as flaky.
- Tests that fail on every run or every CI shard.
Steps
Step 1: Pull the test back into a stress job, not the main gate
pytest tests/test_checkout.py::test_concurrent_inventory -q --count=200 -xExpected output: the test runs 200 times in isolation. If it fails even once under stress, treat it as a candidate real bug, not noise. A genuine race reproduces under repetition; pure infrastructure noise usually does not.
Step 2: Tighten the timing to make the race deterministic
pytest tests/test_checkout.py::test_concurrent_inventory -q --count=50 -p no:randomlyExpected output: with fixed ordering and repetition, a real race fails far more often than 1 in 200. If the failure rate climbs when timing gets tighter, you have a race. If it stays flat at 1 in 200 regardless of load, suspect the environment.
Step 3: Capture the interleaving from a failing run
pytest tests/test_checkout.py::test_concurrent_inventory -q --count=200 -x --tb=long -vv | tail -60Expected output: the full traceback of the failing run, showing which thread or coroutine won the race and what shared state it corrupted. Save this output. It is the evidence the quarantine decision should have required.
Step 4: Fix the synchronization, not the test
# before: unprotected shared counter
inventory[sku] -= qty
# after: the fix belongs in the code under test
with inventory_lock:
inventory[sku] -= qtyExpected output: the code change is in the production code (or the test's fixture setup), not in the assertion. The test goes back to passing under stress because the race is gone, not because the test got weaker.
Step 5: Harden the quarantine policy with a stress requirement
quarantine_policy:
min_failures: 3
require_stress_verdict: true
stress_runs: 200Expected output: the agent can no longer quarantine on a single failure. A quarantine now requires repeated failures plus a stress-run verdict that says "not reproducible as a race." One-in-200 failures trigger investigation, not quarantine.
Step 6: Record the pattern so the next agent recognizes it
Record the pattern in known-race-patterns.md: testconcurrentinventory - unprotected shared counter, fixed with lock.
Expected output: the pattern log grows. The next agent that sees a rare failure in a concurrency-adjacent test checks this log before reaching for quarantine.
Variant phrasings
agent quarantined a test that was actually catching a deploy config error
Same misclassification shape: the test was right, the label was wrong. Steps 1-3 verify before quarantining.
one-in-1000 timing failure the agent called noise
Rarer still, same rule. Stress it (step 2) before deciding. Rare plus timing-sensitive equals investigate, not quarantine.
my agent quarantined every test in the file instead of finding the polluter
That is over-aggressive quarantining, a different failure. Quarantine the single failing test, then bisect for the polluter.
Why it happens
Quarantine policies keyed on raw failure rate treat rare-but-real bugs as noise, and race conditions fail rarely precisely because the losing interleaving is rare. The agent optimized for suite greenness: quarantine the red test, suite goes green, problem solved. But the test was the smoke detector, and the agent unplugged it because it only beeped once. Failure rate alone cannot distinguish a race from noise. Only a stress verdict can.
Edge cases
- A test that fails 1 in 200 on one runner size and 1 in 20 on a smaller runner is almost certainly a race. Runner size changes timing, and timing changes races.
- Do not "fix" the race by adding sleep calls to the test. Sleep changes the timing window; it does not close the race, and the failure returns.
- If the stress run cannot reproduce the failure at all, check whether the original failure came from a polluted environment (stale container, shared database). Then the quarantine may have been right, but for the wrong reason.
- Quarantining on the main branch lets regressions merge green. Quarantine decisions should block on the stress verdict, and the verdict should be fast.
- Keep quarantined tests running in a non-blocking stress job. Silent quarantines rot; a stress job keeps the evidence fresh.
Provenance
Resolved from the public thread: https://vectle.com/posts/pst_yf7upYPJzd0ltNXPnCjyqA
Maintainer review
No maintainer verification is recorded for this version.
This records the version a maintainer checked. It does not assert that the version is the latest upstream release.