how to drain and cordon a NotReady node safely
Drains and cordons a NotReady node safely. Use when a node goes NotReady and you need workloads moved off, before node maintenance or replacement, or when a sick node keeps getting scheduled pods. Covers cordon vs drain, PodDisruptionBudgets, and DaemonSet handling. Not for debugging why the node went NotReady, cluster autoscaling, or deleting nodes.
TL;DR
A NotReady node still holds pods that are effectively down, so move them: cordon the node to stop new scheduling, then drain it to evict pods onto healthy nodes. Respect PodDisruptionBudgets (drain will block rather than violate them), skip DaemonSets explicitly, and give stateful pods time to shut down gracefully. Draining is the safe path; deleting the node object is not.
Error / query
how to drain and cordon a NotReady node safelyUse this skill when
- A node shows NotReady and its pods need relocating
- Preparing a node for maintenance, reboot, or replacement
- A flapping node keeps receiving new pods
- You need workloads off a sick node without downtime
Not for this skill when
- You are diagnosing why the node went NotReady (kubelet troubleshooting)
- The node is already deleted from the cluster
- You want to scale the cluster down (autoscaler handles that)
Steps
Step 1: Cordon the node to stop new scheduling
kubectl cordon [node-name]
kubectl get nodes | grep [node-name]Expected: the node shows SchedulingDisabled. Nothing new lands there, but existing pods keep running (or failing) until drained.
Step 2: Check what will block the drain
kubectl get pdb -A -o wide
kubectl get pods --all-namespaces -o wide --field-selector spec.nodeName=[node-name] | grep -v DaemonSetExpected: you see PodDisruptionBudgets that might block eviction and the pods actually on the node. If a PDB allows zero disruptions, drain will wait forever; plan around it.
Step 3: Drain with graceful handling
kubectl drain [node-name] --ignore-daemonsets --delete-emptydir-data --timeout=300sExpected: pods evict one by one and reschedule on healthy nodes. --ignore-daemonsets skips DaemonSet pods (they cannot be rescheduled elsewhere), --delete-emptydir-data allows pods with ephemeral local storage. The 300s timeout bounds each eviction.
Step 4: Verify the node is empty and repair or replace
kubectl get pods --all-namespaces -o wide --field-selector spec.nodeName=[node-name]Expected: only DaemonSet pods remain (or nothing). Now fix the node (restart kubelet, reboot, reimage) or terminate it, then kubectl uncordon only if the node is healthy again.
Variant phrasings
"kubectl drain notready node"
Cordon first (step 1), then drain (step 3). Draining a NotReady node still works; eviction does not need the kubelet healthy.
"move pods off bad node kubernetes"
Drain is the command. Steps 2-3 handle the PDB and DaemonSet wrinkles.
Why it happens
NotReady means the kubelet stopped reporting, but the pod objects remain bound to the node and the scheduler will not move them on its own. Cordon plus drain is the explicit relocation: cordon freezes the node's intake, drain evicts with PDB awareness, and the controllers recreate the pods on healthy nodes.
Edge cases and pitfalls
- Pods with local storage (
emptyDir) are deleted by drain only with--delete-emptydir-data; their data is lost, so back it up if it matters. - Static pods (manifests on the node) cannot be drained; they are managed by the kubelet, not the API.
- If the node never comes back, delete the node object after draining so the cluster stops tracking it.
- Drain on a single-node cluster evicts everything with nowhere to go; it will block on PDBs indefinitely. Plan maintenance windows accordingly.
Provenance
Resolved from the public thread: https://vectle.com/posts/pst5xDfyppAXYG2efGqB8ATw
Maintainer review
No maintainer verification is recorded for this version.
This records the version a maintainer checked. It does not assert that the version is the latest upstream release.