## TL;DR

A NotReady node means the kubelet stopped reporting: SSH to the node (or use a debug pod), check the kubelet service is running, and read its logs for the real error. Common causes are kubelet crashed, disk full, or the node lost contact with the API server. Fix the kubelet and the node flips back to Ready.

## The error

```text
NAME       STATUS     ROLES   AGE   VERSION
worker-2   NotReady   [none]  45d   v1.31.0
```

Events: `NodeNotReady`, pods on the node get evicted after the toleration timeout.

## Use this when

- `kubectl get nodes` shows NotReady
- pods evicted from a specific node
- kubelet wont start after a reboot or upgrade
- one node flaps between Ready and NotReady

## Not for

- DiskPressure or MemoryPressure (separate conditions, separate fixes)
- nodes you cordoned yourself (uncordon them)
- API server down (then every node looks bad)

## Steps

1. Get the node's reported conditions:

```bash
kubectl describe node [node-name]
```

Expected: Conditions section shows `Ready: False` with a reason like `KubeletNotReady`. The message often names the failing check (PLEG, runtime, network).

2. Check if the kubelet process is alive on the node:

```bash
systemctl status kubelet
journalctl -u kubelet --since "30 min ago" | tail -40
```

Expected: active (running) or a clear failure. On managed clusters without SSH, use `kubectl debug node/[node-name] -it --image=busybox` and run `chroot /host systemctl status kubelet`.

3. Read the kubelet logs for the actual error:

```bash
journalctl -u kubelet --since "1 hour ago" | grep -iE "error|fail" | tail -20
```

Expected: the recurring error. Classics: `failed to get container runtime`, `PLEG is not healthy`, `node not found`, certificate errors talking to the API server.

4. Check the container runtime the kubelet depends on:

```bash
systemctl status containerd
crictl ps 2>&1 | head -5
```

Expected: containerd running and responsive. If the runtime is down, the kubelet cant be Ready; fix the runtime first (disk full is the usual killer, see the node disk skill).

5. Verify the node can reach the API server:

Check connectivity from the node to the API server's health endpoint (port 6443) with curl and a 10 second timeout.

Expected: the endpoint answers `ok`. Timeouts mean a network or firewall problem between node and control plane; the kubelet cant report status it cant deliver.

6. Restart the kubelet after fixing the cause:

```bash
systemctl restart kubelet
kubectl get node [node-name] -w
```

Expected: node flips to Ready within a minute or two. Pods that were evicted get rescheduled; check `kubectl get pods -A -o wide` for stragglers.

### Variant: node flaps Ready/NotReady every few minutes

Usually resource exhaustion (check Pressure conditions) or an unstable network. Also check for clock skew: TLS to the API server fails if the node clock drifts.

### Variant: NotReady right after a node upgrade or reboot

Kubelet config or version mismatch: the new kubelet may reject the old config file. Compare the config against the version's defaults and check for removed flags.

### Variant: all nodes NotReady at once

Thats the control plane, not the kubelets. Check the API server and etcd before touching any node.

## Why it happens

The kubelet heartbeats its status to the API server every 10 seconds; miss enough heartbeats (default 40s grace, then 5m pod eviction toleration) and the node is marked NotReady. The kubelet stops heartbeating when its process dies, when it cant reach the API, or when a core subsystem (runtime, PLEG, CNI) fails its health checks.

## Edge cases

- Pods with long tolerations stay on a NotReady node; dont assume eviction cleaned it up.
- Static pods keep running on a NotReady node even when the API cant see them.
- Cloud provider node lifecycle controllers may terminate long-NotReady nodes; fix fast or the node (and its local data) is gone.
- After recovery, check for pods stuck in Terminating that were mid-eviction; delete them with grace-period 0 if they hang.

## Provenance

Resolved from the public thread: https://vectle.com/posts/pst_sYgC75eBkWC6D-7RjmvqXA
