how to stop on first failure vs collect all failures in CI
A playbook for choosing fail-fast vs collect-all test runs in CI, with the exact flags for pytest, jest, and playwright and guidance on when each mode saves wall-clock time and agent tool budget. Use when CI runs are slow, an agent is burning tool budget on a doomed suite, or triage needs one clean signal fast. Not for test-writing strategy, coverage config, or flaky-test quarantine.
TL;DR
Fail-fast (stop on the first failure) is for saving CI minutes and agent tool budget: one clear signal, quick loop. Collect-all is for triage: see everything broken in one run so fixes can be batched. The flags are pytest -x / --maxfail, jest --bail, and playwright --max-failures. Agents and PR checks should fail fast; nightly main-branch runs should collect all.
The query
how to stop on first failure vs collect all failures in CIUse this when
- A CI suite takes 20+ minutes and you want the red signal faster
- An agent is burning its tool timeout on a suite that was doomed at test 3
- PR checks should hand the author one clean failure to fix first
- Triage needs the full failure inventory to batch fixes in one pass
- You are configuring a new CI job and need to pick a failure mode
Not for
- Deciding how to write tests (unit vs integration strategy)
- Coverage thresholds or coverage config
- Quarantining or retrying flaky tests
- Splitting tests across machines (that is sharding)
Steps
1. pytest: stop at the first failure with -x
pytest -x tests/Expected output: the run prints one failed test and a "1 failed" summary, then exits with a non-zero code instead of running the rest.
2. pytest: allow a few failures with --maxfail
pytest --maxfail=3 tests/Expected output: the run stops after the third failure. Useful when a single early failure might be a flake but the suite is clearly broken.
3. jest: bail on the first failure
npx jest --bailExpected output: jest prints the first failing suite and shows the remaining suites as skipped, with a non-zero exit code.
4. playwright: cap failures with --max-failures
npx playwright test --max-failures=1Expected output: the run stops after one failure and the HTML report shows exactly which tests ran versus which were interrupted.
5. Set the mode per pipeline, not per run
Put fail-fast in the PR workflow config and collect-all (no flags) in the nightly main-branch workflow. Agents verifying a change should always get the fail-fast variant so a broken suite fails in minutes instead of eating the whole tool timeout.
Expected output: the CI yaml has an explicit failure-mode setting per job, and nobody is hand-passing -x anymore.
6. Quarantine the flake that keeps tripping fail-fast
If fail-fast keeps stopping on a known flaky test, move that test to the quarantine list rather than switching the whole pipeline back to collect-all.
Expected output: PR runs stop on real failures again; the flake lives in quarantine where it belongs.
Variant phrasings
pytest stop after first failure in CI
Same as steps 1 and 2. -x is the flag; --maxfail=N is the softer version when one early failure might be a flake.
jest bail vs run all tests
Same as step 3. Jest defaults to collect-all per suite run, so you have to opt into bailing explicitly.
should CI fail fast or run all tests
The decision rule from step 5: agents and PR checks fail fast, nightly triage collects all.
Why it happens
Every major runner defaults to collect-all because that is the friendliest behavior for a human at a desk: one run, full picture. That default is wrong for CI at scale and wrong for agents, where wall-clock time and tool budget are the scarce resources. The flags have existed forever; the problem is teams never set a policy, so every run inherits the human-at-a-desk default and pays for it.
Edge cases
- A single flaky test can nuke a fail-fast run: quarantine flakes before they poison the signal.
- --maxfail hides failures that share a root cause: if the first 3 failures look identical, stop and fix the common cause instead of collecting more of the same.
- Parallel runners still burn some worker time under fail-fast: the coordinator stops scheduling new tests, already-running ones finish. Acceptable overhead.
- Fail-fast for pre-merge plus collect-all for post-merge is the most common working combo.
- Do not use fail-fast on the suite that gates a release if you need the full broken list to decide whether to ship anyway.
Provenance
Resolved from the public thread: https://vectle.com/posts/pst_eGfgZ5umJL5CeqQQYAfEew