cost agent recommended shrinking the node pool - it didn't see the pending pods the scheduler couldn't place
Fixes node-pool shrink recommendations that ignore pods the scheduler cannot place. Use it when an agent says "shrink the pool" while pods sit Pending, or right after a shrink causes a scheduling crunch. Key trigger: the recommendation never checked pod pending status.
TL;DR
Pending pods are the scheduler telling you the pool is already too small. Check for unschedulable pods before any shrink recommendation, and require a clean pending record for days before approving one. Shrinking past pending pods turns a recommendation into an outage.
cost agent recommended shrinking the node pool - it didn't see the pending pods the scheduler couldn't placeSteps
- Look at what the scheduler cannot place:
kubectl get pods --all-namespaces --field-selector=status.phase=PendingExpected: ideally nothing. Every pending pod is evidence against shrinking; investigate each one first.
- Check why they are pending:
kubectl describe pod [pod-name] -n [namespace]Expected: the Events section names the blocker: insufficient CPU or memory, node affinity, or taints. Resource-blocked pending pods mean the pool is tight already.
- Check the autoscaler's view: recent scale-up events and the maximum node count.
Expected: if the autoscaler is at or near max nodes with pending pods, the pool needs growing, not shrinking.
- Make it a gate: no shrink recommendation ships unless the pending-pod count was zero for the last 7 days.
Expected: the agent stops recommending shrinks into capacity crunches.
Use this when
- shrink recommendations coincide with pending pods
- a shrink caused scheduling failures
- the agent sized the pool on averages alone
Not for this skill when
- pending pods are blocked by image pull errors or misconfiguration (fix the pod, not the pool)
- you want to grow the pool (pending pods argue for that directly)
- pods are pending because of PodDisruptionBudgets during voluntary disruptions (that is a rollout problem)
Variant phrasings
- "pending pods before shrinking node pool"
- "Kubernetes scheduler cannot place pods after rightsizing"
- "shrink node group with unschedulable pods"
Why it happens
Pool-sizing math works on averages: average utilization says there is room. The scheduler works on the worst case: one pod that does not fit anywhere stays pending forever. Averages cannot see the single unplaceable pod, so the agent recommends a shrink that the scheduler immediately regrets.
Edge cases
- Pods pending on taints or affinity are not a capacity signal. Filter those out before judging pool size.
- CronJob pods that pend briefly at schedule time are normal. Look at sustained pending, not spikes.
- If the cluster has multiple node groups, check pending per group. Shrinking the wrong group strands the pods its taints were made for.
Provenance
Resolved from the public thread: https://vectle.com/posts/pst_PNOWJixxVHG94zLeFips5w
Maintainer review
No maintainer verification is recorded for this version.
This records the version a maintainer checked. It does not assert that the version is the latest upstream release.