CrashLoopBackOff on one node only: how to isolate node-level causes
Isolates CrashLoopBackOff that hits pods on only one node. Use when the same workload is healthy on other nodes, restarts cluster on one host, or exit codes vary by node. Covers node disk pressure, bad image layers, kubelet config drift, and hardware faults. Not for cluster-wide crash loops, Pending pods, or image-pull failures.
TL;DR
When one node is the common factor behind CrashLoopBackOff, the problem is the node, not the app. Compare the failing pods' placement against healthy ones, then check node disk pressure, kubelet logs, and whether the image layer on that node is corrupt. Fixing the app will not help if the node is sick; cordon, drain, and repair the node instead.
Error / query
CrashLoopBackOff on one node only: how to isolate node-level causesUse this skill when
- The same deployment crash-loops on one node but runs fine on others
- Multiple unrelated pods crash-loop on the same host
- Exit codes or error messages differ by node for the same image
- A crash loop started right after a node upgrade or reboot
Not for this skill when
- Pods crash-loop on every node (fix the app, image, or config)
- Pods are Pending or ImagePullBackOff (different failure class)
- The node shows NotReady (start with kubelet troubleshooting)
Steps
Step 1: Confirm the node is the common factor
kubectl get pods -A -o wide | grep -i crash
kubectl get pods -n [namespace] -o wide | grep [node-name]Expected: every crashing pod lands on the same node while identical pods elsewhere are Running. If crashing pods are spread across nodes, stop here; it is an app problem.
Step 2: Check node conditions and pressure
kubectl describe node [node-name] | grep -A20 ConditionsExpected: DiskPressure, MemoryPressure, or PIDPressure true, or recent KubeletHasDiskPressure events. Pressure explains crashes (eviction, failed writes, OOM) that look like app bugs.
Step 3: Read kubelet and container-runtime logs on the node
journalctl -u kubelet --since "2 hours ago" | tail -50
journalctl -u containerd --since "2 hours ago" | grep -i "error\|fail" | tail -20Expected: errors about pulling layers, writing to disk, or starting containers. A corrupt image layer shows up here even when kubectl describe pod blames the app.
Step 4: Rule out a corrupt image layer on that node
crictl images | grep [image-name]Expected: the image digest on the bad node. If the layer is suspect, delete the image on that node (crictl rmi [image]) and let the next pod re-pull it cleanly. If the re-pulled image still crashes, the node is not the issue.
Step 5: Cordon, drain, and repair the node
kubectl cordon [node-name]
kubectl drain [node-name] --ignore-daemonsets --delete-emptydir-dataExpected: the node stops taking new pods and drains existing ones. Repair (free disk, restart kubelet/containerd, reimage if needed), uncordon, and watch new pods land healthy.
Variant phrasings
"pods crashloop on one node but not others"
Same as above. Step 1 confirms the pattern; do not redeploy the app until the node is ruled out.
"node-specific crashloopbackoff"
Check disk pressure first; it is the most common node-level cause and the cheapest to verify.
Why it happens
Kubernetes schedules identical pods onto different hosts, so a fault that lives on one host (full disk, dying SSD, corrupted pulled layer, kubelet misconfiguration, bad kernel after upgrade) produces the same workload crashing on that host only. The pod events still blame the container, which misleads you into debugging the app.
Edge cases and pitfalls
- Disk pressure can clear itself after evictions, hiding the evidence; check
kubectl describe nodequickly and look at historical events. - A node that just joined after an OS upgrade may have a different containerd or runc version; compare with
kubectl get nodes -o wideand node labels. - Taints on the bad node can concentrate only certain workloads there, making it look node-specific when it is really workload-specific; check tolerations.
- Deleting the image forces a re-pull across the registry; in air-gapped clusters have the image available locally first.
Provenance
Resolved from the public thread: https://vectle.com/posts/pstOldYNm9Q-jfaZtX2PzhdQ
Maintainer review
No maintainer verification is recorded for this version.
This records the version a maintainer checked. It does not assert that the version is the latest upstream release.