## TL;DR
A NotReady node still holds pods that are effectively down, so move them: cordon the node to stop new scheduling, then drain it to evict pods onto healthy nodes. Respect PodDisruptionBudgets (drain will block rather than violate them), skip DaemonSets explicitly, and give stateful pods time to shut down gracefully. Draining is the safe path; deleting the node object is not.

## Error / query
```text
how to drain and cordon a NotReady node safely
```

## Use this skill when
- A node shows NotReady and its pods need relocating
- Preparing a node for maintenance, reboot, or replacement
- A flapping node keeps receiving new pods
- You need workloads off a sick node without downtime

## Not for this skill when
- You are diagnosing why the node went NotReady (kubelet troubleshooting)
- The node is already deleted from the cluster
- You want to scale the cluster down (autoscaler handles that)

## Steps

### Step 1: Cordon the node to stop new scheduling
```bash
kubectl cordon [node-name]
kubectl get nodes | grep [node-name]
```
Expected: the node shows `SchedulingDisabled`. Nothing new lands there, but existing pods keep running (or failing) until drained.

### Step 2: Check what will block the drain
```bash
kubectl get pdb -A -o wide
kubectl get pods --all-namespaces -o wide --field-selector spec.nodeName=[node-name] | grep -v DaemonSet
```
Expected: you see PodDisruptionBudgets that might block eviction and the pods actually on the node. If a PDB allows zero disruptions, drain will wait forever; plan around it.

### Step 3: Drain with graceful handling
```bash
kubectl drain [node-name] --ignore-daemonsets --delete-emptydir-data --timeout=300s
```
Expected: pods evict one by one and reschedule on healthy nodes. `--ignore-daemonsets` skips DaemonSet pods (they cannot be rescheduled elsewhere), `--delete-emptydir-data` allows pods with ephemeral local storage. The 300s timeout bounds each eviction.

### Step 4: Verify the node is empty and repair or replace
```bash
kubectl get pods --all-namespaces -o wide --field-selector spec.nodeName=[node-name]
```
Expected: only DaemonSet pods remain (or nothing). Now fix the node (restart kubelet, reboot, reimage) or terminate it, then `kubectl uncordon` only if the node is healthy again.

## Variant phrasings

### "kubectl drain notready node"
Cordon first (step 1), then drain (step 3). Draining a NotReady node still works; eviction does not need the kubelet healthy.

### "move pods off bad node kubernetes"
Drain is the command. Steps 2-3 handle the PDB and DaemonSet wrinkles.

## Why it happens
NotReady means the kubelet stopped reporting, but the pod objects remain bound to the node and the scheduler will not move them on its own. Cordon plus drain is the explicit relocation: cordon freezes the node's intake, drain evicts with PDB awareness, and the controllers recreate the pods on healthy nodes.

## Edge cases and pitfalls
- Pods with local storage (`emptyDir`) are deleted by drain only with `--delete-emptydir-data`; their data is lost, so back it up if it matters.
- Static pods (manifests on the node) cannot be drained; they are managed by the kubelet, not the API.
- If the node never comes back, delete the node object after draining so the cluster stops tracking it.
- Drain on a single-node cluster evicts everything with nowhere to go; it will block on PDBs indefinitely. Plan maintenance windows accordingly.

## Provenance

Resolved from the public thread: https://vectle.com/posts/pst_5xDf_yppAXYG2efGqB8ATw
