## TL;DR

A pod stuck in Pending means the scheduler couldnt place it, and `kubectl describe pod` tells you why under Events. The usual causes are not enough CPU or memory, node taints, or affinity rules nobody remembers. Read the FailedScheduling message, fix the constraint, and the pod schedules.

## The error

```text
STATUS: Pending
Events:
  0/3 nodes are available: 3 Insufficient memory
```

Variants: `0/2 nodes are available: 1 node(s) had taint`, `pod has unbound immediate PersistentVolumeClaims`, `1 Insufficient cpu`.

## Use this when

- `kubectl get pods` shows STATUS Pending for a long time
- events contain FailedScheduling or `0/N nodes are available`
- a new deploy never produces running pods
- you need to tell resource shortage apart from taint/affinity/PVC causes

## Not for

- CrashLoopBackOff or ImagePullBackOff (pod scheduled, container failing)
- pods stuck in ContainerCreating (image pull or volume mount problems)
- manually suspended jobs or scaled-to-zero deployments (working as intended)

## Steps

1. Read the scheduler's own words:

```bash
kubectl describe pod [pod-name] -n [namespace]
```

Expected: the Events section at the bottom shows `FailedScheduling` with the exact reason, e.g. `0/3 nodes are available: 3 Insufficient memory`. That one line names the fix.

2. If it says Insufficient cpu/memory, check what the pod asks for vs what nodes have:

```bash
kubectl describe node | grep -A 3 "Allocated resources"
kubectl get pods -A -o wide | grep [node-name]
```

Expected: Allocated resources shows how much of each node is already spoken for. Compare the pod's requests against the leftover.

3. Fix resource shortage. Options, in order of pain:

```bash
kubectl top nodes
```

Lower the pod's requests to what it actually uses, scale the cluster up (add nodes or raise the node pool size), or delete dead weight (completed jobs, idle test pods). If limits exceed requests wildly, tighten them so the scheduler sees honest numbers.

4. If it says taint, match the pod to the taint:

```bash
kubectl describe node [node-name] | grep Taints
```

Expected: something like `dedicated=special:NoSchedule`. Fix: add a matching toleration to the pod spec, or remove the taint from the node if it shouldnt be there.

5. If it says node affinity or selector, check the labels line up:

```bash
kubectl get nodes --show-labels | grep [label-key]
```

Expected: at least one node carries the label the pod requires. If no node matches, fix the pod's nodeSelector/affinity or label the nodes correctly.

6. If it says unbound PersistentVolumeClaims, check the PVC:

```bash
kubectl get pvc -n [namespace]
```

Expected: PVC should be Bound. If Pending, the storage class may lack a provisioner or the requested size exceeds what the provisioner allows. Fix the PVC first, the pod follows.

7. Re-check the pod schedules:

```bash
kubectl get pod [pod-name] -n [namespace] -w
```

Expected: STATUS moves Pending to Running within a minute or two of the fix. If it stays Pending, re-run step 1; the event message updates with the next blocker.

### Variant: `0/N nodes are available: N node(s) had untolerated taint`

The whole cluster got tainted (common after upgrades add new taints). Add the toleration to the workload, dont taint-strip every node blindly.

### Variant: Pending only on one specific node pool

Check the pool's labels, taints, and available capacity separately. Autoscaled pools can also be at max size, look for ClusterAutoscaler events.

### Variant: pod Pending with no FailedScheduling events at all

The scheduler might be down or overloaded. Check `kubectl get pods -n kube-system` for the scheduler pods, and look at scheduler logs.

## Why it happens

Scheduling is a filter pass: the scheduler scores every node against the pod's requests, taints, affinity, and volume claims. Pending just means every node failed at least one filter. The FailedScheduling event names the filter that failed, so the fix is almost always literal.

## Edge cases

- ResourceQuota or LimitRange in the namespace can reject the pod before scheduling; check `kubectl describe quota -n [namespace]`.
- PodDisruptionBudgets and topology spread constraints can block scheduling in ways that look like resource shortage.
- GPU or special-hardware pods need the right node labels AND device plugins; Pending with `Insufficient nvidia.com/gpu` means no GPU nodes are free.
- After a cluster upgrade, default taints or admission plugins change; Pending pods that used to schedule fine point at the upgrade, not the workload.

## Provenance

Resolved from the public thread: https://vectle.com/posts/pst__EjEcPpJ-6OmR0L72tF9kQ
