Kubernetes node "MemoryPressure"/"DiskPressure": response checklist
Responds to Kubernetes node MemoryPressure and DiskPressure conditions: identifying the consuming workloads, relieving the pressure, and preventing repeats. Use it when node conditions show pressure or pods go unschedulable citing it. Not for pod-level OOMKilled or full persistent volumes.
Node MemoryPressure / DiskPressure response checklist
TL;DR
Describe the node, find what is eating memory or disk, relieve it (evict the hog, clean the disk), then uncordon the node. For memory the fix is workload limits and eviction; for disk it is image cache, logs, and emptyDir sprawl. The kubelet sets these conditions when usage crosses its thresholds, and the scheduler stops placing pods on the node until they clear.
Kubernetes node "MemoryPressure"/"DiskPressure": response checklistSteps
- Read the node state. Run
kubectl describe node [node]and check the Conditions section plus allocated resources.
Expected: MemoryPressure or DiskPressure shows True, and you see what is allocated vs capacity.
- Find the hogs. Run
kubectl top nodeandkubectl top pod --all-namespaces --sort-by=memory(or cpu) to rank consumers.
Expected: you can name the top memory or disk consumers.
- Relieve memory pressure. Look for pods without memory limits first; evict or delete the offender, or cordon and drain the node if it needs breathing room.
Expected: memory usage drops and the condition flips back to False.
- Relieve disk pressure. Check image disk usage, container log sizes, and kubelet garbage collection. Prune unused images if your runtime and policy allow it.
Expected: disk usage drops below the eviction threshold.
- Return the node to service. Run
kubectl uncordon [node].
Expected: the condition clears and the scheduler places pods on the node again.
- Prevent the repeat. Put memory limits on all workloads, configure log rotation, and size nodes with real headroom.
Expected: no repeat next traffic spike.
Use this when
- Node conditions show MemoryPressure or DiskPressure True
- Pods stay Pending citing node pressure
- You get eviction storms on a node
Not for this skill when
- A single pod was OOMKilled (that is a pod-level memory limit problem)
- A PersistentVolumeClaim is full (different storage, different fix)
- The cluster autoscaler is misbehaving (that is scaling config)
Variant phrasings
- kubernetes node disk pressure
- node memory pressure pods pending
- kubelet eviction thresholds
- node has disk pressure kubernetes
Why it happens
The kubelet watches node-level memory and disk usage. Past its eviction thresholds it marks the pressure condition, starts evicting pods to reclaim the resource, and the scheduler avoids the node. Pressure is the kubelet protecting the node itself at your pods' expense.
Edge cases
- Pressure can flap when usage hovers right at the threshold. Fix the headroom, not the flapping.
- System daemons consume disk too. It is not always your pods.
- A big image pull on a nearly-full disk is enough to tip it over. Watch disk during rollouts.
- Draining a single-node cluster needs care with daemonsets and local data. Plan it, do not improvise it.
Provenance
Resolved from the public thread: https://vectle.com/posts/pst_3VK833eQW6PkbAA1XWyFMg
Maintainer review
No maintainer verification is recorded for this version.
This records the version a maintainer checked. It does not assert that the version is the latest upstream release.