## TL;DR
When one node is the common factor behind CrashLoopBackOff, the problem is the node, not the app. Compare the failing pods' placement against healthy ones, then check node disk pressure, kubelet logs, and whether the image layer on that node is corrupt. Fixing the app will not help if the node is sick; cordon, drain, and repair the node instead.

## Error / query
```text
CrashLoopBackOff on one node only: how to isolate node-level causes
```

## Use this skill when
- The same deployment crash-loops on one node but runs fine on others
- Multiple unrelated pods crash-loop on the same host
- Exit codes or error messages differ by node for the same image
- A crash loop started right after a node upgrade or reboot

## Not for this skill when
- Pods crash-loop on every node (fix the app, image, or config)
- Pods are Pending or ImagePullBackOff (different failure class)
- The node shows NotReady (start with kubelet troubleshooting)

## Steps

### Step 1: Confirm the node is the common factor
```bash
kubectl get pods -A -o wide | grep -i crash
kubectl get pods -n [namespace] -o wide | grep [node-name]
```
Expected: every crashing pod lands on the same node while identical pods elsewhere are Running. If crashing pods are spread across nodes, stop here; it is an app problem.

### Step 2: Check node conditions and pressure
```bash
kubectl describe node [node-name] | grep -A20 Conditions
```
Expected: `DiskPressure`, `MemoryPressure`, or `PIDPressure` true, or recent `KubeletHasDiskPressure` events. Pressure explains crashes (eviction, failed writes, OOM) that look like app bugs.

### Step 3: Read kubelet and container-runtime logs on the node
```bash
journalctl -u kubelet --since "2 hours ago" | tail -50
journalctl -u containerd --since "2 hours ago" | grep -i "error\|fail" | tail -20
```
Expected: errors about pulling layers, writing to disk, or starting containers. A corrupt image layer shows up here even when `kubectl describe pod` blames the app.

### Step 4: Rule out a corrupt image layer on that node
```bash
crictl images | grep [image-name]
```
Expected: the image digest on the bad node. If the layer is suspect, delete the image on that node (`crictl rmi [image]`) and let the next pod re-pull it cleanly. If the re-pulled image still crashes, the node is not the issue.

### Step 5: Cordon, drain, and repair the node
```bash
kubectl cordon [node-name]
kubectl drain [node-name] --ignore-daemonsets --delete-emptydir-data
```
Expected: the node stops taking new pods and drains existing ones. Repair (free disk, restart kubelet/containerd, reimage if needed), uncordon, and watch new pods land healthy.

## Variant phrasings

### "pods crashloop on one node but not others"
Same as above. Step 1 confirms the pattern; do not redeploy the app until the node is ruled out.

### "node-specific crashloopbackoff"
Check disk pressure first; it is the most common node-level cause and the cheapest to verify.

## Why it happens
Kubernetes schedules identical pods onto different hosts, so a fault that lives on one host (full disk, dying SSD, corrupted pulled layer, kubelet misconfiguration, bad kernel after upgrade) produces the same workload crashing on that host only. The pod events still blame the container, which misleads you into debugging the app.

## Edge cases and pitfalls
- Disk pressure can clear itself after evictions, hiding the evidence; check `kubectl describe node` quickly and look at historical events.
- A node that just joined after an OS upgrade may have a different containerd or runc version; compare with `kubectl get nodes -o wide` and node labels.
- Taints on the bad node can concentrate only certain workloads there, making it look node-specific when it is really workload-specific; check tolerations.
- Deleting the image forces a re-pull across the registry; in air-gapped clusters have the image available locally first.

## Provenance

Resolved from the public thread: https://vectle.com/posts/pst_OldYNm9Q-jfa_ZtX2PzhdQ
