Kubernetes node NotReady: kubelet troubleshooting
Troubleshoots Kubernetes nodes stuck in NotReady by checking kubelet health. Use when a node shows NotReady, pods get evicted from it, or kubelet is down or misconfigured. Triggers: node NotReady, kubelet stopped, KubeletNotReady, node lost, pods evicted after node failure. Not for: node resource pressure (DiskPressure/MemoryPressure are separate conditions), control plane outages, intentional cordons and drains.
TL;DR
A NotReady node means the kubelet stopped reporting: SSH to the node (or use a debug pod), check the kubelet service is running, and read its logs for the real error. Common causes are kubelet crashed, disk full, or the node lost contact with the API server. Fix the kubelet and the node flips back to Ready.
The error
NAME STATUS ROLES AGE VERSION
worker-2 NotReady [none] 45d v1.31.0Events: NodeNotReady, pods on the node get evicted after the toleration timeout.
Use this when
kubectl get nodesshows NotReady- pods evicted from a specific node
- kubelet wont start after a reboot or upgrade
- one node flaps between Ready and NotReady
Not for
- DiskPressure or MemoryPressure (separate conditions, separate fixes)
- nodes you cordoned yourself (uncordon them)
- API server down (then every node looks bad)
Steps
- Get the node's reported conditions:
kubectl describe node [node-name]Expected: Conditions section shows Ready: False with a reason like KubeletNotReady. The message often names the failing check (PLEG, runtime, network).
- Check if the kubelet process is alive on the node:
systemctl status kubelet
journalctl -u kubelet --since "30 min ago" | tail -40Expected: active (running) or a clear failure. On managed clusters without SSH, use kubectl debug node/[node-name] -it --image=busybox and run chroot /host systemctl status kubelet.
- Read the kubelet logs for the actual error:
journalctl -u kubelet --since "1 hour ago" | grep -iE "error|fail" | tail -20Expected: the recurring error. Classics: failed to get container runtime, PLEG is not healthy, node not found, certificate errors talking to the API server.
- Check the container runtime the kubelet depends on:
systemctl status containerd
crictl ps 2>&1 | head -5Expected: containerd running and responsive. If the runtime is down, the kubelet cant be Ready; fix the runtime first (disk full is the usual killer, see the node disk skill).
- Verify the node can reach the API server:
Check connectivity from the node to the API server's health endpoint (port 6443) with curl and a 10 second timeout.
Expected: the endpoint answers ok. Timeouts mean a network or firewall problem between node and control plane; the kubelet cant report status it cant deliver.
- Restart the kubelet after fixing the cause:
systemctl restart kubelet
kubectl get node [node-name] -wExpected: node flips to Ready within a minute or two. Pods that were evicted get rescheduled; check kubectl get pods -A -o wide for stragglers.
Variant: node flaps Ready/NotReady every few minutes
Usually resource exhaustion (check Pressure conditions) or an unstable network. Also check for clock skew: TLS to the API server fails if the node clock drifts.
Variant: NotReady right after a node upgrade or reboot
Kubelet config or version mismatch: the new kubelet may reject the old config file. Compare the config against the version's defaults and check for removed flags.
Variant: all nodes NotReady at once
Thats the control plane, not the kubelets. Check the API server and etcd before touching any node.
Why it happens
The kubelet heartbeats its status to the API server every 10 seconds; miss enough heartbeats (default 40s grace, then 5m pod eviction toleration) and the node is marked NotReady. The kubelet stops heartbeating when its process dies, when it cant reach the API, or when a core subsystem (runtime, PLEG, CNI) fails its health checks.
Edge cases
- Pods with long tolerations stay on a NotReady node; dont assume eviction cleaned it up.
- Static pods keep running on a NotReady node even when the API cant see them.
- Cloud provider node lifecycle controllers may terminate long-NotReady nodes; fix fast or the node (and its local data) is gone.
- After recovery, check for pods stuck in Terminating that were mid-eviction; delete them with grace-period 0 if they hang.
Provenance
Resolved from the public thread: https://vectle.com/posts/pst_sYgC75eBkWC6D-7RjmvqXA
Maintainer review
No maintainer verification is recorded for this version.
This records the version a maintainer checked. It does not assert that the version is the latest upstream release.