my agent said "can't reproduce" for a test that fails only when the CI runner's disk is 90% full
Helps an agent investigate CI-only test failures caused by a nearly full runner disk. Use it when a test fails in CI with disk-space, ENOSPC, or write errors but passes locally, and the agent keeps saying it cannot reproduce the failure. Not for genuine logic bugs, and not for failures that also happen on machines with plenty of free disk.
TL;DR The test is not flaky - the CI runner's disk is 90 percent full and the test fails when it cannot write. Free the disk or grow the runner and the "unreproducible" failure goes away.
my agent said "can't reproduce" for a test that fails only when the CI runner's disk is 90% fullSteps
- Stop trying to reproduce it locally. Instead, add a disk check to the CI job: run
df -hearly in the workflow and print free space before the test step.
Expected: the log shows the runner at 90 percent or more used, often on the volume that holds the workspace or the docker data dir.
- Confirm the failure signature matches disk pressure. Look for ENOSPC, "no space left on device", sqlite "database is locked" under parallel writes, or browsers failing to download in the test log.
Expected: at least one error in the log is a write failure, not an assertion failure.
- Find what is eating the disk. Common culprits: old docker layers and images, unpruned build caches, previous job workspaces that were never cleaned, and test artifacts from earlier runs.
Expected: docker system df or a disk-usage scan names one big offender.
- Fix it at the source: add a cleanup step before tests (prune docker, clear caches), split artifacts to external storage, or move the job to a larger runner.
Expected: the next run starts with healthy free space, e.g. under 70 percent used.
- Rerun the previously failing test with nothing else changed.
Expected: green, which proves it was environment, not code.
Use this when
- the agent says "cannot reproduce" but the failure is CI-only
- logs show write errors, ENOSPC, or downloads failing partway
- failures started suddenly across many unrelated tests
- the runner is shared or long-lived rather than ephemeral
Not for this skill when
- the failure also happens on machines with plenty of free disk
- the error is a genuine assertion failure with no write errors nearby
- local and CI both fail identically
- the disk is full because the test itself writes unbounded data (fix the test, not the runner)
Variant phrasings
- "test fails in CI with no space left on device but passes locally"
- "CI runner disk full causing test failures"
- "ENOSPC in CI test run, cannot reproduce locally"
Why it happens
CI runners accumulate state: docker images, layer caches, old checkouts, artifact zips. At around 90 percent full, writes start failing intermittently - a test that needs temp space, a browser download, or a DB file gets a partial write and fails in a way that looks like a flaky assertion. The agent keeps rerunning the test instead of looking at the machine, because the test is the only thing it can see.
Edge cases
- Disk pressure can also slow I/O enough to trip timeouts without any explicit write error; check
dfeven when the error looks like a timeout. - Container jobs share the host disk; the job's own filesystem can look fine while the host is full.
- Some caches refill between runs, so a manual cleanup "fixes" it once and it returns next week; automate the prune step.
Provenance
Resolved from the public thread: https://vectle.com/posts/pst_3L1AdmyEPLrJi-Yy6szwCA
Maintainer review
No maintainer verification is recorded for this version.
This records the version a maintainer checked. It does not assert that the version is the latest upstream release.