## TL;DR

Rare flakes need evidence, not reproduction. Add detailed logging and capture artifacts (traces, videos, dumps) on every run, then wait for the next failure with full forensics ready.

## Error

```text
(Not an error; a debugging strategy. The situation: a test fails weekly and nobody can reproduce it.)
```

## Steps

1. Instrument: enable traces, videos, and verbose logs for that test specifically. Expected: the next failure arrives with evidence.
2. Log the environment with each run: timestamps, runner, load, versions. Expected: correlates failures with conditions.
3. Wait for 2-3 failures and compare the artifacts. Expected: a pattern emerges (time of day, runner type, preceding test).
4. Form one hypothesis from the pattern and test it (for example, pin the timezone). Expected: a targeted fix.
5. Verify over a month of runs. Expected: the flake is gone, proven by absence.

## When to use

- Flakes too rare to reproduce locally.
- You have CI history but no cause.

## When not to use

- Frequent flakes (bisect them directly).
- You already have a hypothesis (test it).

## Tool compatibility

- Any framework; per-test artifact capture.

## Variant phrasings

### Rare flaky test debugging

The same strategy; evidence over reproduction.

### Intermittent test failure investigation

The general case; instrument first.

## Why it happens

Rare flakes depend on rare conditions: specific timing, specific runners, specific preceding tests. Only captured evidence reveals them.

## Edge cases

- Do not change five things at once; you will never know what fixed it.
- Keep the instrumentation after the fix for a month; rare flakes return.
- If it never fails again after instrumentation, the observer effect is real; keep watching.

## Provenance

Resolved from the public thread: https://vectle.com/posts/pst_iv3xE2la-EhEr0GM9k7lsA
