how to measure test suite reliability over time
Describes how to measure test suite reliability over time with rolling pass rates, flake rates, and time-to-green. Use when you want a reliability dashboard or SLA for your suite. Not for diagnosing one broken build.
TL;DR
Measure suite reliability with four numbers tracked on rolling windows: pass rate, flake rate (passed only after retry), mean time to green, and per-test reliability for the worst offenders. Collect per-test outcomes from every CI run into a small database, compute 7-day and 30-day rolling values, and alert when a metric drifts past a threshold. A single number like "the suite is 97 percent reliable" is far more useful than vibes.
The query
how to measure test suite reliability over timeUse this when
- You want a reliability dashboard or a team SLA.
- You need to show whether reliability is improving or degrading.
Not for
- Debugging a single red build.
- Code coverage; that measures something else.
Steps
- Define the metrics: pass rate, flake rate, mean time to green, worst-test reliability. Expected output: written definitions everyone agrees on.
- Store per-test, per-build outcomes (name, build, result, duration, retries) somewhere queryable. Expected output: a growing history table.
- Compute 7-day and 30-day rolling values and chart them. Expected output: trend lines, not just snapshots.
- Set alert thresholds, e.g. flake rate above 5 percent pages the test owner. Expected output: alerts fire on real drift, not noise.
- Review the dashboard weekly and retire metrics nobody looks at. Expected output: a living dashboard the team actually uses.
Provenance
Resolved from the public thread: https://vectle.com/posts/pst_nNeRXzGO3w060AK39q-ePw
Maintainer review
No maintainer verification is recorded for this version.
This records the version a maintainer checked. It does not assert that the version is the latest upstream release.