## TL;DR
Flaky tests fail intermittently without code changes, and they rot trust in CI. The fix is a process, not a patch: detect flakes with rerun data, quarantine them into a non-blocking suite so they stop blocking merges, then triage each one to either fix the root cause or delete the test. Never just rerun the pipeline and hope.

## Error / query
```text
CI flaky tests: how to quarantine and triage
```

## Use this skill when
- The same test passes and fails on identical code across runs
- Developers rerun CI "until it goes green" as a habit
- You need a policy for which tests block merges and which do not
- Flaky failures are the top cause of red builds on main

## Not for this skill when
- A test fails consistently (that is a real bug or a broken test; fix or update it directly)
- The whole suite is slow but stable (performance problem, not flakiness)
- Tests fail only on one developer's machine (environment drift, not CI flakiness)
- You are choosing a test framework (tooling decision, not a flake process)

## Steps

### Step 1: Identify the flakes with data, not anecdotes
```bash
echo "Pull the last 30 runs of the test suite from your CI API and count per-test pass/fail."
echo "A test that fails 5-30% of runs on unchanged code is flaky; near-100% failure is a real break."
```
Expected: a ranked list of flaky tests with failure rates. Anything you cannot measure you cannot triage, so this list is the input to everything below.

### Step 2: Quarantine the flakes out of the blocking suite
```bash
echo "Move flaky tests to a quarantine suite/directory (e.g. tests/quarantine/) that runs but does not gate merges."
echo "Tag them explicitly, e.g. @pytest.mark.quarantine with -m 'not quarantine' on the blocking run."
```
Expected: main stays green, merges unblock, and the flaky tests still run on a schedule so you keep signal on them. Quarantine is containment, not deletion.

### Step 3: Triage each quarantined test into fix, rewrite, or delete
```bash
echo "For each test, classify: timing-dependent (fix with proper waits), order-dependent (isolate state),"
echo "environment-dependent (pin resources), or low-value (delete). Assign an owner and a deadline."
```
Expected: every quarantined test has a disposition and an owner. Tests with no owner and no deadline live in quarantine forever, which is how you got here.

### Step 4: Fix the top root-cause patterns
```bash
echo "Replace sleep() with condition waits, randomize or isolate ports, seed RNGs, and reset shared state in setup/teardown."
```
Expected: the most common flake causes (arbitrary sleeps, port collisions, leaked state between tests, time-zone or clock dependence) are eliminated at the source rather than papered over with retries.

### Step 5: Set a quarantine SLA and enforce it
```bash
echo "Policy: a test may sit in quarantine at most 2 weeks. After that it is fixed, rewritten, or deleted."
echo "Track quarantine size as a metric; review it weekly like any other reliability number."
```
Expected: quarantine stays small and temporary. Without the SLA it becomes a graveyard and developers stop trusting the suite anyway.

## Variant phrasings

### "test passes locally but fails in ci sometimes"
Timing or resource difference. Quarantine it, then reproduce under CI-like constraints (limited CPU, parallel workers) to find the race.

### "how to stop flaky tests blocking PRs"
Quarantine into a non-blocking suite immediately (step 2), then work the triage list. Unblocking merges is the first win.

### "should I retry flaky tests automatically"
Retries hide the signal and slow the suite. Quarantine first; use retries only as a temporary bridge with the flake still tracked for a fix.

## Why it happens
Most flakes are tests that depend on something nondeterministic: timing (sleeps instead of waits), shared mutable state, parallel execution races, external services, or resource contention that only appears under CI load. Each flake trains developers to ignore red builds, which is how real regressions start slipping through.

## Edge cases and pitfalls
- Deleting a flaky test loses coverage; prefer fixing, and only delete when the test's value is genuinely lower than its cost.
- Quarantine suites that nobody watches become invisible; schedule them and alert on new failures.
- A test that is flaky only under `-n auto` parallelism has a shared-state bug; running it serially masks the root cause.
- Do not let "the test is flaky" become the default excuse for real failures; require the failure-rate data from step 1 before quarantining.
- Flakes cluster around infrastructure (docker startup, service readiness); fix the harness once instead of every test.

## Provenance

Resolved from the public thread: https://vectle.com/posts/pst_1GQmhyzNKhdn1Udq5Sl2QA
