## TL;DR

A test whose runtime triples on every failure is not flaky. It is a regression that sometimes crosses the timeout. The agent must compare failure durations against the test's own history before applying the flake label. Add a timing check to the triage policy: if failing runs are consistently 2x or more the passing median, label it a performance regression and open a bug, not a quarantine.

## The query

```text
agent classified a timeout as flaky even though the test's runtime tripled in every failure - how to catch a regression hiding in flake labels
```

## Use this when

- Timeouts are auto-labeled flaky but failing runs are consistently slower than passing runs.
- A test's runtime has drifted upward and the classifier never noticed.
- You suspect a latency regression is hiding behind the flake label.

## Not for

- Genuine network-timeout flakes where failing runs show normal durations.
- Tests with no timing history to compare against.
- Timeouts caused purely by an overloaded CI runner (all tests slow, not one).

## Steps

### Step 1: Pull the timing history for the test

```bash
grep -h "test_payment_retry" reports/junit-*.xml | grep -o 'time="[0-9.]*"' | sort -n
```

Expected output: a sorted list of durations across recent runs. You need at least 20 passing runs to establish a baseline. If the CI system stores timing in a dashboard instead, export the same series from there.

### Step 2: Compare failing-run durations against the passing baseline

```bash
python -c "
import statistics
passing = [4.1, 4.3, 3.9, 4.2, 4.0]
failing = [12.8, 13.1, 12.4]
print('passing median', statistics.median(passing))
print('failing median', statistics.median(failing))
print('ratio', statistics.median(failing) / statistics.median(passing))
"
```

Expected output: the ratio. In the example it is about 3x. Any consistent ratio at or above 2x means the failures are slow for a reason, and "flaky" is the wrong label.

### Step 3: Find what got slow in the failing runs

```bash
grep -B 5 -A 20 "test_payment_retry" ci-logs/failing-run.log | grep -iE "slow|wait|lock|retry|timeout" | head -20
```

Expected output: the slow step inside the test (a database call, an HTTP retry, a lock wait). The tripling runtime has a cause, and it is usually one call that got slower, not the whole test.

### Step 4: Relabel and file the regression

Record the decision in triage-decisions.log: test_payment_retry - 3x runtime on failures, relabeled perf-regression, not flaky.

Expected output: the decision is recorded with the evidence (the ratio from step 2). The test is now tracked as a performance regression with an owner, not silently quarantined as a flake.

### Step 5: Fix the classifier to require timing evidence

```yaml
flake_classifier:
  require_timing_check: true
  max_fail_to_pass_ratio: 2.0
```

Expected output: the agent can no longer label a timeout flaky without comparing durations. A timeout with a 2x-plus slowdown ratio is routed to the regression queue automatically.

### Step 6: Add a runtime-drift alert for the suite

Add a monitoring rule: alert if any test's 7-day median runtime exceeds 1.5x its 30-day median.

Expected output: the next slowdown gets caught as drift before it starts tripping timeouts, so the classifier never has to choose between "flaky" and "slow" again.

## Variant phrasings

### agent bumped the timeout to 60s and called it fixed
Bumping the timeout hides the regression the way quarantine hides the race. Steps 2-3 find the real slowdown; the timeout stays where it was.

### flake classifier trusts green reruns but the rerun ran on a different runner
A green rerun on a bigger runner proves nothing about the code. Step 2's ratio uses the test's own history on comparable runners.

### test passes on retry so the agent closed it as flaky
Passing on retry is necessary but not sufficient evidence. The timing check in step 5 is the missing half of the verdict.

## Why it happens

Timeouts have two causes: slow code and slow environments. The classifier only checked the symptom (it timed out, then passed on retry) and never the durations. A regression that triples runtime will pass on retry whenever the slow path happens to be fast that run, which looks exactly like flakiness to a classifier that ignores timing. The label was wrong because the evidence was incomplete, not because the classifier is stupid.

## Edge cases

- If all tests on the runner are slow during the failing runs, the cause is the runner, not the test. Compare the test's slowdown against the suite-wide slowdown before relabeling.
- A test with high natural variance (integration tests hitting real services) needs a wider ratio threshold. Tune step 5's 2.0 per test type.
- Do not delete the timeout to "prove" it is slow. The timeout is the safety net. Measure around it.
- The first few runs after a deploy can be slow (cold caches, JIT warmup). Exclude post-deploy runs from the baseline or the ratio will false-positive.
- If the slowdown is in test setup rather than the code under test, the regression is in the test infrastructure. File it there, not against the product code.

## Provenance

Resolved from the public thread: https://vectle.com/posts/pst_rdLRUM5-ibp_J7c7KYozYw
