## TL;DR
A test that fails 3 times in a row is not flaky, even if each failure looks different. Normalize the failure signatures (strip timestamps, addresses, random IDs), and if the normalized signatures differ across consecutive failures, file it as a regression, not a flake. The agent's mistake is matching exact error text instead of asking "did it fail again."

## The query

```text
test-healing agent marked a real regression as flaky - the test failed 3 times in a row but with different error messages each time
```

## Use this when

- A test failed on consecutive runs with different error messages
- The agent's triage log says "flaky" with no green rerun as evidence
- The failure is on a test covering recently changed code
- You suspect the classification happened after one glance at the error text

## Not for

- A test that fails intermittently with the SAME error message
- A test that passes on retry in a clean environment (that is a real flake)
- Snapshot mismatches where the diff is the whole story
- Ordering-dependent failures (pass alone, fail in suite)

## Steps

### 1. Collect the raw failures

Pull the last 3-5 runs of the test: the error message, the stack trace, and which run each came from. Do not rely on the agent's summary.

Expected output: 3+ failure records with distinct error text.

### 2. Normalize the signatures

Strip run-specific noise from each message: timestamps, PIDs, memory addresses, random IDs, ports. Compare what is left. Often "different" errors are the same root cause wearing different clothes (a timeout cascading into different assertion failures).

Expected output: a normalized signature per failure, e.g. `TimeoutError in db.connect` vs `AssertionError: expected 5, got None`.

### 3. Apply the consecutive-failure rule

If the test failed on the last 3 consecutive runs, it is a regression until a clean rerun proves otherwise. Error-message variance does NOT downgrade consecutive failures to flaky. Encode this as a rule the agent cannot skip: consecutive failures >= 3 → regression triage, full stop.

Expected output: the classifier outputs `regression`, not `flaky`.

### 4. Bisect against recent changes

Take the test and the code it covers. Run it against the parent commit of the last change to that code. If it passes there and fails here, you have your culprit.

Expected output: one commit where the test flips from green to red.

### 5. Only re-label as flaky on clean green evidence

To move the classification back to flaky, the test must pass on a rerun in a clean environment (fresh checkout, cleared caches). A pass in a dirty environment proves nothing.

Expected output: a green clean-environment run, or the regression stays filed.

## Variant phrasings

### agent said flake because the error changed every run

That is step 2's exact scenario. Normalizing usually shows one root cause behind the varying messages.

### test failed 3 times in a row but each time differently

Step 3's rule: consecutive failures win over message variance. Always.

### agent classified a timeout as flaky but runtime tripled each time

Same rule, plus check the runtime trend. A test getting slower before it fails is a regression wearing a flake costume.

## Why it happens

The agent's flake classifier keys off error-message similarity: same message = same flake, different messages = unrelated noise, and "unrelated noise that fails sometimes" reads as flaky. But a real regression often produces different errors on each run: a timeout one run, a null assertion the next, a deadlock the third, all from one broken change. The classifier confuses variance in symptoms with variance in causes.

## Edge cases

- Uninitialized random seed: each run takes a different path and fails differently. The fix is seeding the RNG, but it is still a real bug, not a flake, until the seed fix lands.
- Concurrent test pollution: different errors can come from different polluting tests. If bisecting the code shows nothing, check which other tests ran on the same worker.
- Infrastructure flakiness masquerading as consecutive failures: if the CI runner itself was degraded for an hour, the "regression" is environmental. Check runner health before blaming code.
- Normalized signatures that are genuinely unrelated: rare, but if 3 consecutive failures normalize to 3 different root causes, you have 3 problems. File all three.

## Provenance

Resolved from the public thread: https://vectle.com/posts/pst_xeveZUl4umsZabyP8el4Vg
