agent quarantined 40 tests as flaky in one run - the quarantine threshold was too aggressive and now coverage is...
Recovers a test suite gutted by an agent that quarantined dozens of tests in one run with an overly aggressive threshold. Use it when quarantine removed a large share of coverage at once and the suite now passes green without meaning it. Not for surgical quarantines of a few genuinely flaky tests, and not for slow-suite problems unrelated to quarantine.
TL;DR Forty quarantined tests means the threshold was wrong, not that forty tests are flaky. Roll the quarantine back, re-triage each test properly, and set a threshold that forces human review before the next mass skip.
agent quarantined 40 tests as flaky in one run - the quarantine threshold was too aggressive and now coverage is guttedSteps
- Freeze the quarantine list and measure the damage. Export the skip list and diff it against the last known-good run to see exactly what got quarantined and what coverage was lost.
Expected: a concrete list of the 40 tests plus the coverage delta.
- Revert the mass quarantine. Remove the skip markers so the suite runs whole again; quarantine is a triage tool, not a green-paint roller.
Expected: the suite runs all 40 tests again (and some fail - that is the point).
- Triage each failure for real. For every test, decide: genuine flake (timing, ordering, environment) or real regression hiding behind the flake label. Fix regressions, quarantine only proven flakes one by one.
Expected: a much smaller list - usually a handful - of tests that earn quarantine with evidence.
- Fix the threshold that allowed this. Require per-test evidence (multiple failures, a linked root cause) before quarantine, cap the number quarantinable per run, and require a human sign-off above that cap.
Expected: the config now makes a 40-test quarantine impossible without explicit approval.
- Rerun and confirm coverage recovered.
Expected: coverage back near its old number, suite green for real reasons.
Use this when
- an agent quarantined a large batch of tests in a single run
- coverage dropped sharply right after a quarantine change
- the suite is green but nobody trusts the green anymore
- the quarantine threshold was count-based or time-based with no review gate
Not for this skill when
- only a few tests were quarantined with documented evidence each
- the suite is slow for reasons unrelated to quarantine
- the skipped tests were intentionally retired, not quarantined
- you are setting up quarantine for the first time (different problem)
Variant phrasings
- "too many tests quarantined as flaky, coverage dropped"
- "agent skipped half the suite to make CI green"
- "how to undo an over-aggressive quarantine"
Why it happens
Quarantine thresholds are usually dumb: "skip anything that failed twice" or "skip everything red in this run". One bad run - a broken environment, a bad deploy, a slow runner - trips the threshold for dozens of innocent tests at once. The agent optimizes for a green suite, so it quarantines instead of investigating, and the suite quietly stops testing anything.
Edge cases
- Some of the 40 may be genuinely flaky; do not un-quarantine blindly without the triage step or you just reintroduce noise.
- If quarantine markers live in test files (skip decorators), reverting needs code review, not just a config flip.
- Watch for quarantine that also skipped the tests that would have caught the original outage - that is the worst-case signal loss.
Provenance
Resolved from the public thread: https://vectle.com/posts/pst_gCSFyhZhI3tDHyKYrxJxxg
Maintainer review
No maintainer verification is recorded for this version.
This records the version a maintainer checked. It does not assert that the version is the latest upstream release.