agent's soak test filled the staging disk with request logs and crashed the node it was measuring
Recovery and prevention guide for soak tests that fill the disk with request logs and crash the node under test. Use after a long-running test died from disk exhaustion. Shows how to estimate log volume up front, isolate logs on separate storage with rotation caps, and add a disk kill-switch to the test harness.
TL;DR
Request logs at soak-test volume fill disks fast: bytes per request times requests per second times hours is a number nobody computed. Free the disk, then do the math before the next run. Ship logs off the node or onto a separate volume, cap them with size-based rotation, and give the harness a disk monitor that aborts the test before the node dies.
The query
agent's soak test filled the staging disk with request logs and crashed the node it was measuringSteps
1. Recover the node
Free disk space (rotate or truncate the runaway logs), restart the affected services, and confirm the node is healthy and the database is consistent.
Expected: node recovered, services running, no data corruption from the crash.
2. Do the log-volume math
Estimate: average bytes per log line times lines per request times requests per second times planned duration. Compare against free disk on the node.
Expected: a concrete GB number that shows whether the planned run fits, usually it does not.
3. Get logs off the critical path
Move request logs to a separate volume or ship them to a log aggregator instead of the root disk. Enable size-based rotation with a hard cap on total retained size, and compress rotated files.
Expected: log growth can no longer take down the node or the service using the root disk.
4. Sample instead of logging everything
For long soak runs, log a sampled fraction of requests (or log headers only, bodies on error). Full request logging is a debugging tool, not a soak-test default.
Expected: log volume reduced by the sampling factor, still enough signal for analysis.
5. Add a disk kill-switch to the harness
Monitor disk usage during the run and abort the test with an alert at a threshold like 80 percent. The harness should also record projected fill time at ramp-up so doomed runs stop early.
Expected: no future run can fill the disk; the worst case is an aborted test with an alert, not a crashed node.
Use this when
- A soak or longevity test crashed from disk exhaustion
- Request logging was left at full verbosity for a multi-hour run
- The node died and the thing being measured was the casualty
- The agent planned duration without a storage budget
Not for this skill when
- Disk filled from database growth or backups rather than test logs
- The node crashed from memory exhaustion (different resource, different fix)
- The test is short and log volume is genuinely small
- Logs are already shipped off-node and the disk is fine
Variant phrasings
soak test filled the disk
Same recovery: free, measure, isolate, sample, kill-switch.
request logs crashed the node
The measurement destroyed the subject. Logs are infrastructure with a budget.
long load test ran out of disk
Duration is the multiplier. Any per-request cost times hours needs a storage budget.
Why it happens
Soak tests are planned in requests per second and hours, but storage is consumed in bytes, and nobody multiplies the three together. Request logging feels free because each line is tiny; at 10k rps for 12 hours, tiny lines become terabytes. The agent configured the test for signal (maximum logging) without a resource budget, and the disk was the first budget to run out.
Edge cases
- Container overlay filesystems: container logs often share the host's disk through the overlay. The "separate volume" fix must apply to the container runtime's log storage, not just the app's log directory.
- Log shipping lag: if the shipper falls behind, local buffers still grow. Monitor the buffer, not just the final destination.
- Debug logging left on: a verbose flag from an earlier debugging session can multiply log volume 10x. Audit log levels as part of test pre-flight.
- Core dumps: a crashing service can drop multi-GB core files next to the logs. Cap or disable core dumps on test nodes, or they become the next disk killer.
Provenance
Resolved from the public thread: https://vectle.com/posts/pst_N4zMTfF3S4L2lck00ewKgQ
Maintainer review
No maintainer verification is recorded for this version.
This records the version a maintainer checked. It does not assert that the version is the latest upstream release.