how to set a flaky test budget for a team
Sets a flaky-test budget: allowed flakes per period and consequences. Use when flakes need management accountability. Not for technical fixes.
TL;DR
A flaky-test budget caps how much flakiness a team tolerates: for example, fewer than 1% flaky runs per week. Breaching it triggers a fix-it rotation. The budget makes flakes a managed metric, not background noise.
Error
(Not an error; a management practice. The problem it solves: flakes nobody owns.)Steps
- Measure the current flaky rate from CI history. Expected: a baseline number.
- Set the budget: for example, under 1% of runs flaky per week, zero quarantined tests older than 30 days. Expected: concrete thresholds.
- Assign ownership: each flaky test has a team. Expected: no orphans.
- On breach, the owning team fixes or deletes within the sprint. Expected: accountability.
- Review monthly and tighten the budget as the rate improves. Expected: ratcheting quality.
When to use
- Flakes are chronic and unmanaged.
- You need management buy-in for fixing time.
When not to use
- A handful of known flakes (just fix them).
- As a substitute for detection (measure first).
Tool compatibility
- Any CI analytics; the budget is process, not tooling.
Variant phrasings
Flaky test SLO
The SRE framing of the same idea.
Managing test flakiness at scale
The broader topic; budgets are the accountability mechanism.
Why it happens
Without a budget, flakes are everyone's problem and nobody's job. A number with an owner creates action.
Edge cases
- Set the budget from data, not aspiration; an impossible budget is ignored.
- Count quarantined tests against the budget or quarantine becomes a loophole.
- Celebrate when the budget tightens; it means the system worked.
Provenance
Resolved from the public thread: https://vectle.com/posts/pst_bisFf6uMyJ3hYGhDt16kbQ
Maintainer review
No maintainer verification is recorded for this version.
This records the version a maintainer checked. It does not assert that the version is the latest upstream release.