my agent kept re-running a failing test locally where it passed - the CI runner had 2GB RAM and the test needed 4
Fixes agents that burn local reruns while the real problem is CI's RAM ceiling. Use when a test fails only in CI, local runs pass, and the signature smells like memory pressure (exit 137, OOM kills). Confirm the kill in CI logs, then raise the runner's memory or slim the test's footprint. Not for failures with no memory signal, and not for accepted heavy-by-design tests.
TL;DR
Stop rerunning it locally; your machine has the RAM and CI does not. Check the CI logs for the OOM killer (exit code 137 is the tell), then either give the job a bigger runner or slim the test's memory footprint. The reruns were never going to reproduce a memory ceiling they do not have.
The exact query
my agent kept re-running a failing test locally where it passed - the CI runner had 2GB RAM and the test needed 4Steps
- Stop the local reruns and read the CI logs for memory evidence: look for exit code 137 (OOM kill), "Killed" messages, GC thrashing warnings, or the job dying at the same memory-heavy step every time. One 137 in the log ends the debate.
Expected: A concrete memory signal from CI, for example "container exited 137 during the fixture that loads the 3GB dataset".
- Measure the test's actual footprint: run it locally under memory observation (container with a 2GB limit, or /usr/bin/time -v, or the language's memory profiler) and record the peak. Compare against the CI runner's limit.
Expected: Numbers, for example "peak 3.8GB vs CI limit 2GB". The gap is the diagnosis.
- Choose the fix deliberately: if the test's memory use is legitimate (large dataset, heavy browser), raise the CI job's memory to fit it and document the requirement next to the test. If the footprint is accidental (loading the whole dataset for a test that needs 10 rows, a leak across tests), slim the test instead.
Expected: Either the runner is sized to the test with the requirement documented, or the test is slimmed to fit the runner. No silent mismatch remains.
- Apply the fix and verify under the real limit: run the test in a container capped at the CI memory limit. It must pass with headroom, not exactly at the edge.
Expected: Green under the cap with margin. A test that passes at 1.99GB on a 2GB runner will fail the next time anything shifts.
- Add a memory guardrail: a CI step that reports peak memory per test (or per job), and an alert when a test's footprint crosses a threshold of the runner's limit. Future memory pressure shows up as a trend, not a mystery failure.
Expected: The next test that outgrows its runner is caught by the trend line before it starts failing.
Use this when
- A test fails only in CI and always passes locally
- CI logs show exit 137, "Killed", or the job dying at a memory-heavy step
- The CI runner has much less RAM than the machines used for reproduction
- An agent keeps rerunning locally instead of checking CI's memory ceiling
Not for this skill when
- There is no memory signal in the failure (no 137, no kill, plenty of headroom measured)
- The test is heavy by accepted design and the runner was already budgeted for it (then something else regressed; check what changed)
- The failure is a hard error (assertion, exception) rather than a death or slowdown
Variant phrasings
CI container exited with code 137 during tests
That is the OOM killer. Find the memory hog, size the runner or slim the test.
test passes locally but gets killed in CI
Memory ceiling difference. Measure the footprint, compare to the CI limit.
how to tell if CI test failure is OOM
Exit code 137, "Killed" in the log, death at the same heavy step every time. Then measure peak vs the runner limit.
Why it happens
Local machines are generous (16 to 64GB) and CI runners are not (2 to 4GB is common on free tiers). A test that casually uses 4GB never notices the difference locally, but in CI the OOM killer notices. The agent reproduces where the constraint does not exist, so every rerun passes and the investigation stalls. Memory pressure also degrades gracefully before it kills: GC thrashing makes tests slow and flaky before the 137 arrives, which sends the agent chasing timing ghosts.
Edge cases
- Memory leaks across tests can make the killer strike a test that is not itself heavy; the footprint measurement must cover the whole worker process, not just the single test in isolation.
- Parallel runners multiply memory: 4 workers each using 1GB need 4GB plus overhead. Size for workers times per-test peak, not a single test.
- Some languages over-commit or the container limit differs from the CI-advertised limit. Measure inside the actual CI container image, not just on your laptop with a limit flag.
- Swap can mask the problem locally (slow but alive) while CI has no swap (dead). Disable swap in your local reproduction for fidelity.
Provenance
Resolved from the public thread: https://vectle.com/posts/pstWLc9NDLFYzCJFPZ9GFrVQ