Kubernetes HPA not scaling: metrics-server debugging
Debugs Kubernetes HPA not scaling by checking metrics-server health and metric flow. Use when HPA shows unknown metrics, replicas never change under load, or metrics-server is missing or broken. Triggers: HPA unable to get metrics, metrics not available, metrics-server CrashLoopBackOff, kubectl top empty. Not for: HPA scaling too aggressively (tune thresholds), custom metrics adapters (Prometheus adapter), Cluster Autoscaler node issues.
TL;DR
HPA cant scale without metrics, and metrics come from metrics-server: check the HPA status for unknown metrics, verify metrics-server is running and kubectl top works, then look at the HPA's target vs current values. Nine times out of ten the metrics pipeline is broken, not the HPA config.
The error
NAME REFERENCE TARGETS MINPODS MAXPODS REPLICAS
myapp Deploy/myapp [unknown]/70% 2 10 2Also: failed to get cpu utilization: unable to get metrics for resource cpu.
Use this when
- HPA shows
[unknown]in TARGETS - replicas never move under obvious load
kubectl top podsreturns nothing or errors- metrics-server pods are crashlooping or missing
Not for
- HPA scaling too fast or flapping (threshold and stabilization tuning)
- custom or external metrics (thats the prometheus adapter, different setup)
- nodes not scaling (Cluster Autoscaler, different component)
Steps
- Read what the HPA thinks:
kubectl describe hpa [hpa-name] -n [namespace]Expected: Conditions explain the state. ScalingActive=False with FailedGetResourceMetric means the metric query failed; the events name the exact error.
- Check metrics-server exists and is healthy:
kubectl get pods -n kube-system -l k8s-app=metrics-server
kubectl top nodesExpected: metrics-server Running, kubectl top returns numbers. If top fails, the pipeline is broken upstream of the HPA.
- If metrics-server is crashlooping, read its logs:
kubectl logs -n kube-system deployment/metrics-server --tail=30Expected: common cause is TLS verification against kubelets (x509: certificate signed by unknown authority). The usual fix is the --kubelet-insecure-tls flag on self-signed clusters, set in the metrics-server deployment args.
- Verify the metric path end to end:
kubectl get --raw /apis/metrics.k8s.io/v1beta1/namespaces/[namespace]/pods | head -c 300Expected: JSON with pod metrics. If this 404s, the metrics API isnt registered; if it errors, metrics-server cant reach kubelets (network policy or firewall on 10250).
- Check the HPA can actually see the target metric:
kubectl get hpa [hpa-name] -n [namespace] -o jsonpath='{.spec.metrics}'Expected: CPU utilization targets need CPU requests set on the containers; without requests, utilization has no denominator and stays unknown. Set requests and the metric starts flowing.
- Generate load and watch it react:
kubectl get hpa [hpa-name] -n [namespace] -wExpected: TARGETS shows 85%/70% style numbers and REPLICAS climbs within a minute or two of sustained load. HPA reacts to averages over its sync period, so give it time.
Variant: metrics work but HPA still shows unknown
The HPA object may predate the metric fix; delete and recreate it, or wait for its next sync. Also check the HPA targets the right workload name.
Variant: scales up but never back down
Thats stabilization windows and scale-down policy, not a metrics problem. Check behavior.scaleDown.stabilizationWindowSeconds.
Variant: metrics-server fine, but one namespace has no metrics
RBAC: the HPA controller needs get on the metrics API. Check for custom RBAC or admission policies blocking it.
Why it happens
HPA is a control loop: read metric, compare to target, adjust replicas. metrics-server feeds it by scraping kubelets. Break any link (server down, TLS to kubelets failing, no CPU requests to compute utilization against) and the HPA parks at unknown and changes nothing. The HPA config is rarely the problem.
Edge cases
- metrics-server scrapes every 60s by default; HPA decisions lag real load by a couple minutes, dont expect instant reaction.
- CPU requests of 0 or unset make utilization undefined; always set requests on HPA-targeted workloads.
- Multiple HPAs on one workload fight each other; keep one HPA per scalable target.
- On managed clusters metrics-server is preinstalled; if someone deleted it, reinstall from the upstream manifests rather than debugging the absence.
Provenance
Resolved from the public thread: https://vectle.com/posts/pst_9uTihaxGd0QH8pkjiHS5Tw
Maintainer review
No maintainer verification is recorded for this version.
This records the version a maintainer checked. It does not assert that the version is the latest upstream release.