## TL;DR

No space left on device on a node usually means old container images and dead container layers piled up, and the fix is pruning them: `crictl rmi --prune` or docker system prune on the node, then find whats actually eating disk with `du`. If it keeps coming back, set kubelet eviction thresholds and image garbage collection so it cleans itself.

## The error

```text
failed to pull image: write /var/lib/containerd/...: no space left on device
Warning  Evicted  The node was low on resource: ephemeral-storage
```

Node condition: `DiskPressure=True` in `kubectl describe node`.

## Use this when

- pods wont start with `no space left on device`
- image pulls fail with write errors
- node shows DiskPressure=True
- pods get evicted for ephemeral-storage

## Not for

- a PVC or mounted volume filling up (thats storage, not the node disk)
- an app writing huge logs inside its container (check log rotation)
- disk full on a VM outside Kubernetes

## Steps

1. Confirm which node and how full:

```bash
kubectl describe node [node-name] | grep -B 2 -A 2 Pressure
df -h /var/lib/containerd
```

Expected: DiskPressure=True, and df shows the container runtime partition at or near 100 percent.

2. Find the biggest consumers on the node:

```bash
du -sh /var/lib/containerd/* 2>/dev/null | sort -rh | head
du -sh /var/log/* 2>/dev/null | sort -rh | head -5
```

Expected: one of images, container layer storage, or /var/log dominates. Old unused images are the usual winner.

3. Prune unused images (safe: only removes images no container references):

```bash
crictl rmi --prune
```

Expected: lists removed images and frees gigabytes. On docker-based nodes use `docker system prune -a -f` instead, but check nothing references the images first.

4. Clean dead containers and their writable layers:

```bash
crictl rm -a -f 2>/dev/null; crictl rmi --prune
```

Expected: stopped containers gone, more space back. Running containers are untouched.

5. If logs are the hog, cap them. Check the kubelet log settings:

```bash
du -sh /var/log/pods/* 2>/dev/null | sort -rh | head -5
```

Expected: one noisy pod's logs dominate. Fix with containerLogMaxSize in the kubelet config (e.g. 50Mi, 5 files) so logs rotate instead of growing forever.

6. Make it self-healing so it doesnt recur. Kubelet config:

```yaml
evictionHard:
  nodefs.available: "10%"
imageGCHighThresholdPercent: 80
imageGCLowThresholdPercent: 60
```

Expected: kubelet starts garbage-collecting images at 80 percent full and evicts disk-hungry pods at 10 percent free, before the node wedges.

7. Verify the node recovers:

```bash
kubectl describe node [node-name] | grep DiskPressure
df -h /var/lib/containerd
```

Expected: DiskPressure=False, usage comfortably under 80 percent, pending pods start scheduling.

### Variant: ephemeral-storage eviction of a specific pod

The pod wrote too much to its writable layer or emptyDir. Give it a real volume, or set ephemeral-storage requests/limits so the scheduler places it on roomy nodes.

### Variant: disk fills again within days

Something is actively writing: usually unrotated logs or a runaway cache. Find the writer with `du` over time before pruning again, or youll prune forever.

### Variant: managed cluster, no SSH to nodes

Use a privileged DaemonSet debug pod or the cloud console's node shell. Same commands, just run through `kubectl debug node/[node-name]`.

## Why it happens

Kubelet image garbage collection is conservative by default and container logs dont rotate unless configured. Over weeks, pulled images, dead layers, and logs accumulate until the disk hits 100 percent, and then everything that writes (image pulls, new pods) fails at once.

## Edge cases

- Pruning images forces re-pulls on next schedule; on slow networks that delays pod startup, so prune during quiet hours.
- `crictl rmi --prune` wont remove images still referenced by any container spec, even stopped ones; remove dead containers first.
- Separate partitions for /var/lib/containerd vs root: check which partition df flags, prune the right one.
- If the disk is 100 percent full you may not even be able to SSH in; use the cloud console serial shell or debug pod as a fallback.

## Provenance

Resolved from the public thread: https://vectle.com/posts/pst_vakiSoskSOd10jCZrDvEGg
