how to measure flake rate per test over time
Measures flake rate per test over time to find unreliable tests. Use when CI is flaky and you need data on which tests flake, prioritizing quarantine candidates, or tracking whether deflaking work helps. Covers collecting pass/fail history, computing flake rates, and dashboards. Not for fixing a specific flaky test, build infra failures, or test authoring.
TL;DR
You cannot fix flakiness you cannot measure: record every test's pass/fail per run, then compute flake rate as runs-with-mixed-results over total runs. Rank tests by flake rate times run frequency; the top of that list is your quarantine backlog. Track the rate weekly; if it is not dropping, your deflaking work is not landing.
Error / query
how to measure flake rate per test over timeUse this skill when
- CI fails intermittently and nobody knows which tests are guilty
- You need a quarantine backlog ordered by evidence
- Tracking whether flakiness is getting better or worse
- Justifying time spent on test reliability
Not for this skill when
- One specific test flakes and you want the root cause (debug that test)
- Builds fail on infra (agents, disks, network)
- Writing new tests (authoring guidance)
Steps
Step 1: Emit machine-readable test results from CI
# example: pytest with junit output
pytest --junitxml=test-results/results.xml
# then upload results.xml as a CI artifact every run, pass or failExpected: every CI run produces a structured result file, including runs that fail. Without results from failed runs you cannot compute rates; make artifact upload unconditional.
Step 2: Store per-test outcomes over time
# minimal schema: (date, suite, test_name, outcome)
# outcome in: passed, failed, skipped, flaky(passed-on-retry)Expected: a table or log you can query by test and time window. A simple append-only file in object storage works; you do not need a database to start.
Step 3: Compute flake rate per test
# flake rate = runs where the test both passed and failed (across retries/reruns)
# / total runs of that test, over the last 14 days
# rank by: flake_rate * runs_per_day (impact ordering)Expected: a ranked list. A test flaking 2 percent of the time but running 500 times a day outranks one flaking 50 percent but running twice; impact ordering focuses effort where CI pain is highest.
Step 4: Dashboard it and act on the ranking
Track weekly: total flake rate, top 10 flaky tests, quarantined count.
Rule: any test above 5 percent flake rate gets quarantined until fixed.Expected: the dashboard shows whether the top offenders change week to week. Quarantine removes the CI pain immediately; the ranking tells you what to fix next.
Variant phrasings
"track flaky tests over time"
Steps 1-3: structured results, stored history, computed rates.
"flaky test dashboard"
Step 4. The metric that matters is flake rate times frequency, trended weekly.
Why it happens
Flakiness is a rate, not a binary property, and humans are bad at estimating rates from anecdotes: the test that failed loudly yesterday feels flakier than the one failing quietly 3 percent of the time for months. Measurement replaces anecdote with a ranked list, which is what turns CI is flaky into actionable work.
Edge cases and pitfalls
- Tests that fail on retry differently than on first run need retry-aware recording; count a test as flaky if any attempt in the run disagreed.
- Infra failures (agent died, disk full) look like test failures; tag them separately or your flake rates blame innocent tests.
- New tests have no history; give them a burn-in window before ranking them against established tests.
- Quarantined tests still run (just not gating); if you stop running them, you lose the signal that the fix worked.
Provenance
Resolved from the public thread: https://vectle.com/posts/pst_L7G1x67SFdzQODS0im08Ug