## TL;DR

HPA cant scale without metrics, and metrics come from metrics-server: check the HPA status for `unknown` metrics, verify metrics-server is running and `kubectl top` works, then look at the HPA's target vs current values. Nine times out of ten the metrics pipeline is broken, not the HPA config.

## The error

```text
NAME   REFERENCE   TARGETS            MINPODS   MAXPODS   REPLICAS
myapp  Deploy/myapp  [unknown]/70%   2         10        2
```

Also: `failed to get cpu utilization: unable to get metrics for resource cpu`.

## Use this when

- HPA shows `[unknown]` in TARGETS
- replicas never move under obvious load
- `kubectl top pods` returns nothing or errors
- metrics-server pods are crashlooping or missing

## Not for

- HPA scaling too fast or flapping (threshold and stabilization tuning)
- custom or external metrics (thats the prometheus adapter, different setup)
- nodes not scaling (Cluster Autoscaler, different component)

## Steps

1. Read what the HPA thinks:

```bash
kubectl describe hpa [hpa-name] -n [namespace]
```

Expected: Conditions explain the state. `ScalingActive=False` with `FailedGetResourceMetric` means the metric query failed; the events name the exact error.

2. Check metrics-server exists and is healthy:

```bash
kubectl get pods -n kube-system -l k8s-app=metrics-server
kubectl top nodes
```

Expected: metrics-server Running, `kubectl top` returns numbers. If top fails, the pipeline is broken upstream of the HPA.

3. If metrics-server is crashlooping, read its logs:

```bash
kubectl logs -n kube-system deployment/metrics-server --tail=30
```

Expected: common cause is TLS verification against kubelets (`x509: certificate signed by unknown authority`). The usual fix is the `--kubelet-insecure-tls` flag on self-signed clusters, set in the metrics-server deployment args.

4. Verify the metric path end to end:

```bash
kubectl get --raw /apis/metrics.k8s.io/v1beta1/namespaces/[namespace]/pods | head -c 300
```

Expected: JSON with pod metrics. If this 404s, the metrics API isnt registered; if it errors, metrics-server cant reach kubelets (network policy or firewall on 10250).

5. Check the HPA can actually see the target metric:

```bash
kubectl get hpa [hpa-name] -n [namespace] -o jsonpath='{.spec.metrics}'
```

Expected: CPU utilization targets need CPU requests set on the containers; without requests, utilization has no denominator and stays unknown. Set requests and the metric starts flowing.

6. Generate load and watch it react:

```bash
kubectl get hpa [hpa-name] -n [namespace] -w
```

Expected: TARGETS shows `85%/70%` style numbers and REPLICAS climbs within a minute or two of sustained load. HPA reacts to averages over its sync period, so give it time.

### Variant: metrics work but HPA still shows unknown

The HPA object may predate the metric fix; delete and recreate it, or wait for its next sync. Also check the HPA targets the right workload name.

### Variant: scales up but never back down

Thats stabilization windows and scale-down policy, not a metrics problem. Check `behavior.scaleDown.stabilizationWindowSeconds`.

### Variant: metrics-server fine, but one namespace has no metrics

RBAC: the HPA controller needs get on the metrics API. Check for custom RBAC or admission policies blocking it.

## Why it happens

HPA is a control loop: read metric, compare to target, adjust replicas. metrics-server feeds it by scraping kubelets. Break any link (server down, TLS to kubelets failing, no CPU requests to compute utilization against) and the HPA parks at unknown and changes nothing. The HPA config is rarely the problem.

## Edge cases

- metrics-server scrapes every 60s by default; HPA decisions lag real load by a couple minutes, dont expect instant reaction.
- CPU requests of 0 or unset make utilization undefined; always set requests on HPA-targeted workloads.
- Multiple HPAs on one workload fight each other; keep one HPA per scalable target.
- On managed clusters metrics-server is preinstalled; if someone deleted it, reinstall from the upstream manifests rather than debugging the absence.

## Provenance

Resolved from the public thread: https://vectle.com/posts/pst_9uTihaxGd0QH8pkjiHS5Tw
