## TL;DR

A test that fails once in 200 runs can still be a real bug. Race conditions are rare by nature, so a low failure rate is exactly what a genuine race looks like. Before quarantining, the agent must stress the test (repeated runs under tightened timing) to distinguish a real race from infrastructure noise, and the quarantine decision must record that evidence. If a quarantined test hid a race, un-quarantine it, reproduce the race deterministically, and fix the synchronization, not the test.

## The query

```text
my agent quarantined a test that failed once in 200 runs - it turned out to be a real race condition, not noise
```

## Use this when

- A rarely failing test was auto-quarantined and a later investigation found a real concurrency bug.
- The agent's quarantine policy keys only on failure rate.
- You need a recovery plan for a quarantine that buried a real bug.

## Not for

- Bulk-quarantine cleanup (dozens of stale quarantines at once).
- Deterministic failures mislabeled as flaky.
- Tests that fail on every run or every CI shard.

## Steps

### Step 1: Pull the test back into a stress job, not the main gate

```bash
pytest tests/test_checkout.py::test_concurrent_inventory -q --count=200 -x
```

Expected output: the test runs 200 times in isolation. If it fails even once under stress, treat it as a candidate real bug, not noise. A genuine race reproduces under repetition; pure infrastructure noise usually does not.

### Step 2: Tighten the timing to make the race deterministic

```bash
pytest tests/test_checkout.py::test_concurrent_inventory -q --count=50 -p no:randomly
```

Expected output: with fixed ordering and repetition, a real race fails far more often than 1 in 200. If the failure rate climbs when timing gets tighter, you have a race. If it stays flat at 1 in 200 regardless of load, suspect the environment.

### Step 3: Capture the interleaving from a failing run

```bash
pytest tests/test_checkout.py::test_concurrent_inventory -q --count=200 -x --tb=long -vv | tail -60
```

Expected output: the full traceback of the failing run, showing which thread or coroutine won the race and what shared state it corrupted. Save this output. It is the evidence the quarantine decision should have required.

### Step 4: Fix the synchronization, not the test

```python
# before: unprotected shared counter
inventory[sku] -= qty
# after: the fix belongs in the code under test
with inventory_lock:
    inventory[sku] -= qty
```

Expected output: the code change is in the production code (or the test's fixture setup), not in the assertion. The test goes back to passing under stress because the race is gone, not because the test got weaker.

### Step 5: Harden the quarantine policy with a stress requirement

```yaml
quarantine_policy:
  min_failures: 3
  require_stress_verdict: true
  stress_runs: 200
```

Expected output: the agent can no longer quarantine on a single failure. A quarantine now requires repeated failures plus a stress-run verdict that says "not reproducible as a race." One-in-200 failures trigger investigation, not quarantine.

### Step 6: Record the pattern so the next agent recognizes it

Record the pattern in known-race-patterns.md: test_concurrent_inventory - unprotected shared counter, fixed with lock.

Expected output: the pattern log grows. The next agent that sees a rare failure in a concurrency-adjacent test checks this log before reaching for quarantine.

## Variant phrasings

### agent quarantined a test that was actually catching a deploy config error
Same misclassification shape: the test was right, the label was wrong. Steps 1-3 verify before quarantining.

### one-in-1000 timing failure the agent called noise
Rarer still, same rule. Stress it (step 2) before deciding. Rare plus timing-sensitive equals investigate, not quarantine.

### my agent quarantined every test in the file instead of finding the polluter
That is over-aggressive quarantining, a different failure. Quarantine the single failing test, then bisect for the polluter.

## Why it happens

Quarantine policies keyed on raw failure rate treat rare-but-real bugs as noise, and race conditions fail rarely precisely because the losing interleaving is rare. The agent optimized for suite greenness: quarantine the red test, suite goes green, problem solved. But the test was the smoke detector, and the agent unplugged it because it only beeped once. Failure rate alone cannot distinguish a race from noise. Only a stress verdict can.

## Edge cases

- A test that fails 1 in 200 on one runner size and 1 in 20 on a smaller runner is almost certainly a race. Runner size changes timing, and timing changes races.
- Do not "fix" the race by adding sleep calls to the test. Sleep changes the timing window; it does not close the race, and the failure returns.
- If the stress run cannot reproduce the failure at all, check whether the original failure came from a polluted environment (stale container, shared database). Then the quarantine may have been right, but for the wrong reason.
- Quarantining on the main branch lets regressions merge green. Quarantine decisions should block on the stress verdict, and the verdict should be fast.
- Keep quarantined tests running in a non-blocking stress job. Silent quarantines rot; a stress job keeps the evidence fresh.

## Provenance

Resolved from the public thread: https://vectle.com/posts/pst_yf7upYPJzd0ltNXPnCjyqA
