CI flaky tests: how to quarantine and triage
Builds a process for handling flaky CI tests. Use when tests pass and fail on identical code, developers rerun pipelines until green, or flakes are the top cause of red builds. Covers flake detection, quarantine suites, triage dispositions, and SLAs. Not for consistently failing tests or slow-but-stable suites.
TL;DR
Flaky tests fail intermittently without code changes, and they rot trust in CI. The fix is a process, not a patch: detect flakes with rerun data, quarantine them into a non-blocking suite so they stop blocking merges, then triage each one to either fix the root cause or delete the test. Never just rerun the pipeline and hope.
Error / query
CI flaky tests: how to quarantine and triageUse this skill when
- The same test passes and fails on identical code across runs
- Developers rerun CI "until it goes green" as a habit
- You need a policy for which tests block merges and which do not
- Flaky failures are the top cause of red builds on main
Not for this skill when
- A test fails consistently (that is a real bug or a broken test; fix or update it directly)
- The whole suite is slow but stable (performance problem, not flakiness)
- Tests fail only on one developer's machine (environment drift, not CI flakiness)
- You are choosing a test framework (tooling decision, not a flake process)
Steps
Step 1: Identify the flakes with data, not anecdotes
echo "Pull the last 30 runs of the test suite from your CI API and count per-test pass/fail."
echo "A test that fails 5-30% of runs on unchanged code is flaky; near-100% failure is a real break."Expected: a ranked list of flaky tests with failure rates. Anything you cannot measure you cannot triage, so this list is the input to everything below.
Step 2: Quarantine the flakes out of the blocking suite
echo "Move flaky tests to a quarantine suite/directory (e.g. tests/quarantine/) that runs but does not gate merges."
echo "Tag them explicitly, e.g. @pytest.mark.quarantine with -m 'not quarantine' on the blocking run."Expected: main stays green, merges unblock, and the flaky tests still run on a schedule so you keep signal on them. Quarantine is containment, not deletion.
Step 3: Triage each quarantined test into fix, rewrite, or delete
echo "For each test, classify: timing-dependent (fix with proper waits), order-dependent (isolate state),"
echo "environment-dependent (pin resources), or low-value (delete). Assign an owner and a deadline."Expected: every quarantined test has a disposition and an owner. Tests with no owner and no deadline live in quarantine forever, which is how you got here.
Step 4: Fix the top root-cause patterns
echo "Replace sleep() with condition waits, randomize or isolate ports, seed RNGs, and reset shared state in setup/teardown."Expected: the most common flake causes (arbitrary sleeps, port collisions, leaked state between tests, time-zone or clock dependence) are eliminated at the source rather than papered over with retries.
Step 5: Set a quarantine SLA and enforce it
echo "Policy: a test may sit in quarantine at most 2 weeks. After that it is fixed, rewritten, or deleted."
echo "Track quarantine size as a metric; review it weekly like any other reliability number."Expected: quarantine stays small and temporary. Without the SLA it becomes a graveyard and developers stop trusting the suite anyway.
Variant phrasings
"test passes locally but fails in ci sometimes"
Timing or resource difference. Quarantine it, then reproduce under CI-like constraints (limited CPU, parallel workers) to find the race.
"how to stop flaky tests blocking PRs"
Quarantine into a non-blocking suite immediately (step 2), then work the triage list. Unblocking merges is the first win.
"should I retry flaky tests automatically"
Retries hide the signal and slow the suite. Quarantine first; use retries only as a temporary bridge with the flake still tracked for a fix.
Why it happens
Most flakes are tests that depend on something nondeterministic: timing (sleeps instead of waits), shared mutable state, parallel execution races, external services, or resource contention that only appears under CI load. Each flake trains developers to ignore red builds, which is how real regressions start slipping through.
Edge cases and pitfalls
- Deleting a flaky test loses coverage; prefer fixing, and only delete when the test's value is genuinely lower than its cost.
- Quarantine suites that nobody watches become invisible; schedule them and alert on new failures.
- A test that is flaky only under
-n autoparallelism has a shared-state bug; running it serially masks the root cause. - Do not let "the test is flaky" become the default excuse for real failures; require the failure-rate data from step 1 before quarantining.
- Flakes cluster around infrastructure (docker startup, service readiness); fix the harness once instead of every test.
Provenance
Resolved from the public thread: https://vectle.com/posts/pst_1GQmhyzNKhdn1Udq5Sl2QA
Maintainer review
No maintainer verification is recorded for this version.
This records the version a maintainer checked. It does not assert that the version is the latest upstream release.