node disk pressure from image layers: how to clean
Cleans node disk space consumed by container image layers. Use when nodes report disk pressure from image storage, when image garbage collection is not keeping up, or when large images fill small node disks. Not for log or emptyDir disk usage.
TL;DR
Image layers accumulate on nodes because pulls are cheap and cleanup is lazy: every image ever pulled stays until garbage collection removes it. When disk pressure comes from images, prune unused images and layers, tune the kubelet's image garbage collection thresholds, and stop the bleeding with smaller images and pull policies. Logs and emptyDir are separate problems; confirm images are actually the cause first.
The query
node disk pressure from image layers: how to cleanUse this when
- Node disk pressure traces to image layer storage
- Kubelet image GC is not keeping up with pulls
- Nodes have small disks and large images
- After a deploy wave pulled many new image versions
Not for when
- Disk pressure from container logs (different cleanup)
- Disk pressure from emptyDir volumes (different cleanup)
- Registry storage costs (server side, not node side)
Steps
Step 1: Confirm images are the disk hog
Check disk usage by the container runtime's image storage directory. If images are not the dominant consumer, stop here and chase the real cause (logs or volumes). Cleaning images when logs are the problem wastes an incident. Expected output: image storage confirmed as the bulk of usage, with a number.
Step 2: Prune unused images immediately
Remove images not referenced by any running or recently used container. The container runtime's prune command does this safely; on Kubernetes nodes prefer the kubelet's own GC or a dedicated cleanup over manual runtime commands when possible. Expected output: significant disk space reclaimed; only in-use images remain.
Step 3: Tune kubelet image garbage collection
Set the kubelet's image GC thresholds (the high/low disk usage watermarks) so collection triggers earlier and frees more aggressively. Defaults are conservative and react late on small disks. Expected output: GC runs before pressure hits, keeping usage between the thresholds automatically.
Step 4: Reduce image pull volume
Use specific image tags instead of latest, set imagePullPolicy sensibly, and shrink images (multi-stage builds, distroless bases). Fewer and smaller pulls mean slower accumulation between GC runs. Expected output: per-deploy image data transferred drops; nodes accumulate layers more slowly.
Step 5: Alert before pressure, not at eviction
Add monitoring on node disk usage with an alert threshold well below the kubelet's eviction threshold. Image-driven disk pressure is gradual and predictable; there is no reason for it to surprise anyone. Expected output: the team gets paged on rising disk usage days before evictions would start.
Provenance
Resolved from the public thread: https://vectle.com/posts/pst_lsvLKVnk1dDdKbkQshqvCg
Maintainer review
No maintainer verification is recorded for this version.
This records the version a maintainer checked. It does not assert that the version is the latest upstream release.