## TL;DR
Disk pressure evictions mean the node's disk filled up, and Kubernetes started evicting pods to protect the node. Recover in two phases: get workloads running elsewhere now (cordon the node, let pods reschedule), then find what ate the disk (usually logs, image layers, or emptyDir volumes) and clean it up. If you skip phase two, the evictions come back.

## The query
```text
pod evicted due to node disk pressure: how to recover
```

## Use this when
- Pods show Evicted with disk pressure as the reason
- A node reports DiskPressure condition
- Evictions recur on the same node or nodes
- You need workloads back quickly after an eviction storm

## Not for when
- Memory pressure evictions (different resource, different fix)
- Pods evicted by the descheduler or PDB violations
- Disk issues on persistent volumes (node disk vs PV are different)

## Steps

### Step 1: Cordon the pressured node and let workloads move
Mark the node unschedulable so new pods do not land on a full disk. Evicted pods with controllers (Deployments, StatefulSets) reschedule automatically; check that they come up healthy elsewhere.
Expected output: workloads running on healthy nodes; the pressured node holds no new pods.

### Step 2: Find what filled the disk
Check the big three in order: container log files under the kubelet log directory, unused image layers, and emptyDir volumes that grew unbounded. A quick disk usage sort on the node's filesystem usually names the culprit in seconds.
Expected output: one directory or volume type accounting for most of the usage.

### Step 3: Clean up safely
Remove unused images and stopped containers, rotate or truncate runaway log files, and check for emptyDir volumes without size limits. Do not delete live container logs for running pods you are still debugging; copy what you need first.
Expected output: disk usage back under the eviction thresholds; the DiskPressure condition clears.

### Step 4: Add emptyDir size limits and log rotation
Set sizeLimit on emptyDir volumes so one pod cannot fill a node, and confirm container log rotation is configured (max size and max files). These two settings prevent the most common recurrences.
Expected output: the same workload pattern can no longer fill a node disk.

### Step 5: Uncordon and set up alerting
Once disk pressure clears and stays clear, uncordon the node. Add an alert on node disk usage well before the eviction threshold so the next fill-up pages someone instead of evicting pods.
Expected output: node schedulable again, with an early-warning alert in place.

## Provenance

Resolved from the public thread: https://vectle.com/posts/pst_DNJPMsMgQB-8OBp2Q97TfQ
