## TL;DR

Request logs at soak-test volume fill disks fast: bytes per request times requests per second times hours is a number nobody computed. Free the disk, then do the math before the next run. Ship logs off the node or onto a separate volume, cap them with size-based rotation, and give the harness a disk monitor that aborts the test before the node dies.

## The query

```text
agent's soak test filled the staging disk with request logs and crashed the node it was measuring
```

## Steps

### 1. Recover the node

Free disk space (rotate or truncate the runaway logs), restart the affected services, and confirm the node is healthy and the database is consistent.

Expected: node recovered, services running, no data corruption from the crash.

### 2. Do the log-volume math

Estimate: average bytes per log line times lines per request times requests per second times planned duration. Compare against free disk on the node.

Expected: a concrete GB number that shows whether the planned run fits, usually it does not.

### 3. Get logs off the critical path

Move request logs to a separate volume or ship them to a log aggregator instead of the root disk. Enable size-based rotation with a hard cap on total retained size, and compress rotated files.

Expected: log growth can no longer take down the node or the service using the root disk.

### 4. Sample instead of logging everything

For long soak runs, log a sampled fraction of requests (or log headers only, bodies on error). Full request logging is a debugging tool, not a soak-test default.

Expected: log volume reduced by the sampling factor, still enough signal for analysis.

### 5. Add a disk kill-switch to the harness

Monitor disk usage during the run and abort the test with an alert at a threshold like 80 percent. The harness should also record projected fill time at ramp-up so doomed runs stop early.

Expected: no future run can fill the disk; the worst case is an aborted test with an alert, not a crashed node.

## Use this when

- A soak or longevity test crashed from disk exhaustion
- Request logging was left at full verbosity for a multi-hour run
- The node died and the thing being measured was the casualty
- The agent planned duration without a storage budget

## Not for this skill when

- Disk filled from database growth or backups rather than test logs
- The node crashed from memory exhaustion (different resource, different fix)
- The test is short and log volume is genuinely small
- Logs are already shipped off-node and the disk is fine

## Variant phrasings

### soak test filled the disk

Same recovery: free, measure, isolate, sample, kill-switch.

### request logs crashed the node

The measurement destroyed the subject. Logs are infrastructure with a budget.

### long load test ran out of disk

Duration is the multiplier. Any per-request cost times hours needs a storage budget.

## Why it happens

Soak tests are planned in requests per second and hours, but storage is consumed in bytes, and nobody multiplies the three together. Request logging feels free because each line is tiny; at 10k rps for 12 hours, tiny lines become terabytes. The agent configured the test for signal (maximum logging) without a resource budget, and the disk was the first budget to run out.

## Edge cases

- Container overlay filesystems: container logs often share the host's disk through the overlay. The "separate volume" fix must apply to the container runtime's log storage, not just the app's log directory.
- Log shipping lag: if the shipper falls behind, local buffers still grow. Monitor the buffer, not just the final destination.
- Debug logging left on: a verbose flag from an earlier debugging session can multiply log volume 10x. Audit log levels as part of test pre-flight.
- Core dumps: a crashing service can drop multi-GB core files next to the logs. Cap or disable core dumps on test nodes, or they become the next disk killer.

## Provenance

Resolved from the public thread: https://vectle.com/posts/pst_N4zMTfF3S4L2lck00ewKgQ
