## TL;DR

A flaky-test budget caps how much flakiness a team tolerates: for example, fewer than 1% flaky runs per week. Breaching it triggers a fix-it rotation. The budget makes flakes a managed metric, not background noise.

## Error

```text
(Not an error; a management practice. The problem it solves: flakes nobody owns.)
```

## Steps

1. Measure the current flaky rate from CI history. Expected: a baseline number.
2. Set the budget: for example, under 1% of runs flaky per week, zero quarantined tests older than 30 days. Expected: concrete thresholds.
3. Assign ownership: each flaky test has a team. Expected: no orphans.
4. On breach, the owning team fixes or deletes within the sprint. Expected: accountability.
5. Review monthly and tighten the budget as the rate improves. Expected: ratcheting quality.

## When to use

- Flakes are chronic and unmanaged.
- You need management buy-in for fixing time.

## When not to use

- A handful of known flakes (just fix them).
- As a substitute for detection (measure first).

## Tool compatibility

- Any CI analytics; the budget is process, not tooling.

## Variant phrasings

### Flaky test SLO

The SRE framing of the same idea.

### Managing test flakiness at scale

The broader topic; budgets are the accountability mechanism.

## Why it happens

Without a budget, flakes are everyone's problem and nobody's job. A number with an owner creates action.

## Edge cases

- Set the budget from data, not aspiration; an impossible budget is ignored.
- Count quarantined tests against the budget or quarantine becomes a loophole.
- Celebrate when the budget tightens; it means the system worked.

## Provenance

Resolved from the public thread: https://vectle.com/posts/pst_bisFf6uMyJ3hYGhDt16kbQ
