## TL;DR
You cannot fix flakiness you cannot measure: record every test's pass/fail per run, then compute flake rate as runs-with-mixed-results over total runs. Rank tests by flake rate times run frequency; the top of that list is your quarantine backlog. Track the rate weekly; if it is not dropping, your deflaking work is not landing.

## Error / query
```text
how to measure flake rate per test over time
```

## Use this skill when
- CI fails intermittently and nobody knows which tests are guilty
- You need a quarantine backlog ordered by evidence
- Tracking whether flakiness is getting better or worse
- Justifying time spent on test reliability

## Not for this skill when
- One specific test flakes and you want the root cause (debug that test)
- Builds fail on infra (agents, disks, network)
- Writing new tests (authoring guidance)

## Steps

### Step 1: Emit machine-readable test results from CI
```bash
# example: pytest with junit output
pytest --junitxml=test-results/results.xml
# then upload results.xml as a CI artifact every run, pass or fail
```
Expected: every CI run produces a structured result file, including runs that fail. Without results from failed runs you cannot compute rates; make artifact upload unconditional.

### Step 2: Store per-test outcomes over time
```bash
# minimal schema: (date, suite, test_name, outcome)
# outcome in: passed, failed, skipped, flaky(passed-on-retry)
```
Expected: a table or log you can query by test and time window. A simple append-only file in object storage works; you do not need a database to start.

### Step 3: Compute flake rate per test
```bash
# flake rate = runs where the test both passed and failed (across retries/reruns)
#           / total runs of that test, over the last 14 days
# rank by: flake_rate * runs_per_day (impact ordering)
```
Expected: a ranked list. A test flaking 2 percent of the time but running 500 times a day outranks one flaking 50 percent but running twice; impact ordering focuses effort where CI pain is highest.

### Step 4: Dashboard it and act on the ranking
```text
Track weekly: total flake rate, top 10 flaky tests, quarantined count.
Rule: any test above 5 percent flake rate gets quarantined until fixed.
```
Expected: the dashboard shows whether the top offenders change week to week. Quarantine removes the CI pain immediately; the ranking tells you what to fix next.

## Variant phrasings

### "track flaky tests over time"
Steps 1-3: structured results, stored history, computed rates.

### "flaky test dashboard"
Step 4. The metric that matters is flake rate times frequency, trended weekly.

## Why it happens
Flakiness is a rate, not a binary property, and humans are bad at estimating rates from anecdotes: the test that failed loudly yesterday feels flakier than the one failing quietly 3 percent of the time for months. Measurement replaces anecdote with a ranked list, which is what turns `CI is flaky` into actionable work.

## Edge cases and pitfalls
- Tests that fail on retry differently than on first run need retry-aware recording; count a test as flaky if any attempt in the run disagreed.
- Infra failures (agent died, disk full) look like test failures; tag them separately or your flake rates blame innocent tests.
- New tests have no history; give them a burn-in window before ranking them against established tests.
- Quarantined tests still run (just not gating); if you stop running them, you lose the signal that the fix worked.

## Provenance

Resolved from the public thread: https://vectle.com/posts/pst_L7G1x67SFdzQODS0im08Ug
