flaky test dashboard: what metrics to track
Defines the metrics for a flaky-test dashboard: rates, age, and ownership. Use when building flake visibility. Not for the fixing workflow.
TL;DR
Track flakiness rate per test, time-in-quarantine, and owner for each flaky test. Those three numbers drive every decision: fix, quarantine, or delete.
Error
(Not an error; a metrics definition. The gap it fills: "we have flakes" with no numbers.)Steps
- Track per test: runs, failures, failure rate over 14 days. Expected: a ranked flaky list.
- Track quarantine age: days since quarantined per test. Expected: stale quarantines visible.
- Track owner: every flaky test has a team. Expected: accountability.
- Track CI impact: minutes lost to retries and reruns. Expected: the cost of flakes in time.
- Review weekly with the ranked list. Expected: the worst offenders get fixed first.
When to use
- Building a flake dashboard or report.
- Justifying flake-fixing time to management.
When not to use
- Fixing a specific flake (use the debugging skills).
- Small suites where the list fits in your head.
Tool compatibility
- Any CI analytics; simple database plus a chart.
Variant phrasings
Flaky test metrics
The general topic; rate, age, owner.
Measure test flakiness
The goal; 14-day failure rate is the core metric.
Why it happens
Without numbers, flakes are anecdotes and never get prioritized. With numbers, they are a ranked backlog.
Edge cases
- Failure rate needs enough runs; weight recent runs more.
- Distinguish infra flakes from test flakes in the metrics.
- Do not rank by count alone; a 50%-flaky test beats ten 1%-flaky ones for attention.
Provenance
Resolved from the public thread: https://vectle.com/posts/pst_53dF1gjuWOInaa5ZKVifKw