## TL;DR

GitHub-hosted runners only have about 14 GB free, and Docker layers plus caches eat it fast. Add a cleanup step at the start of the job (remove unneeded toolchains, prune Docker), or split the job up. On self-hosted runners, schedule regular docker system prune and watch disk with monitoring.

## The error

```text
Error: No space left on device
write /var/lib/docker/...: no space left on device
ENOSPC: no space left on device
```

Jobs fail at checkout, build, or Docker steps, often intermittently.

## Use this when

- Actions jobs fail with no space left on device
- docker build or docker pull fails on the runner
- self-hosted runner disk fills over days
- failures are intermittent (depends what ran before)

## Not for

- Kubernetes node disk pressure
- artifact too large to upload (thats a size limit, not disk)
- your laptop being full

## Steps

1. Confirm disk is the issue and find the hog. Add a debug step:

```yaml
- name: disk usage
  run: df -h && du -sh /opt/* /usr/share/* 2>/dev/null | sort -rh | head
```

Expected: / at or near 100 percent. On GitHub-hosted runners the usual winners are /opt tool caches, Docker layers, and the hosted tool cache.

2. Free the big preinstalled chunks you dont need (github-hosted runners):

```yaml
- name: free disk space
  run: |
    sudo rm -rf /usr/share/dotnet /opt/ghc /usr/local/share/boost
    sudo rm -rf "$AGENT_TOOLSDIRECTORY"
```

Expected: several GB back. Only remove toolchains your build doesnt use; check what your job needs first. There are community actions that do this cleanup, but the manual version is transparent.

3. Prune Docker aggressively if the job uses containers:

```yaml
- name: prune docker
  run: docker system prune -af --volumes
```

Expected: unused images, containers, and build cache gone. Run it BEFORE the docker build step, not after the failure.

4. Shrink what the job stores:

- use `--squash`-style minimal final images or multi-stage builds that drop build layers
- dont cache the world: cache only dependency dirs, and set cache size limits
- upload artifacts instead of keeping everything on disk between jobs

Expected: the job's own footprint drops, so it fits even on a dirty runner.

5. For self-hosted runners, automate the cleanup:

```bash
docker system prune -af --volumes --filter "until=24h"
```

Run this on a cron on every runner host, and alert on disk above 80 percent. Also clear old work dirs: runners keep `_work` from every past job.

6. Re-run and confirm headroom:

```yaml
- name: verify space
  run: df -h / | tail -1
```

Expected: comfortably under 85 percent at the heaviest step. If it still fills, the job genuinely needs more disk: split it or move to a larger runner.

### Variant: fails only on docker build with big contexts

Add a .dockerignore; people ship node_modules and .git into the build context, which multiplies disk use during build.

### Variant: intermittent, passes on re-run

Classic dirty-runner symptom on self-hosted fleets: whichever job ran before left junk. The prune step at job start makes it deterministic.

### Variant: needs more than cleanup can give

GitHub's larger runners have more disk. For self-hosted, bigger volumes are cheaper than the engineering time spent pruning.

## Why it happens

Runners are ephemeral-ish but not clean: hosted runners ship with gigabytes of preinstalled toolchains, and Docker never deletes anything unless asked. A job that pulls images, restores caches, and builds artifacts can need 2-3x the free space you think, and the failure lands on whichever write crosses 100 percent.

## Edge cases

- Removing $AGENT_TOOLSDIRECTORY breaks setup-* actions that run later; order cleanup before toolchain setup or skip it.
- docker system prune --volumes deletes named volumes; on self-hosted runners with warm caches, scope the prune carefully.
- Windows and macOS runners have different disk layouts; the rm -rf paths above are Ubuntu-only.
- Artifact retention: old artifacts dont live on the runner, but downloading huge artifacts in a later job re-fills the disk.

## Provenance

Resolved from the public thread: https://vectle.com/posts/pst_2ZLdAX-VQFBNIzfb4OKSbw
