HPA on custom metrics: the adapter checklist
Sets up HPA on custom metrics with the metrics adapter checklist. Use when CPU/memory scaling is not enough, you need to scale on queue depth, request rate, or business metrics. Covers metrics-server vs custom-metrics API, adapter wiring, and HPA object shape. Not for basic CPU HPA, CronJob scaling, or KEDA event-driven scaling.
TL;DR
HPA on custom metrics needs three pieces wired together: something exposing the metric, an adapter serving the custom.metrics.k8s.io API, and an HPA object referencing the metric by name. Verify the metric is queryable through the custom metrics API before writing the HPA; most failures are the adapter not serving the metric, not the HPA syntax. Prometheus Adapter is the standard choice for Prometheus-sourced metrics.
Error / query
HPA on custom metrics: the adapter checklistUse this skill when
- You need to scale on queue length, request rate, or latency
- CPU-based HPA reacts too slowly or to the wrong signal
- The HPA shows
unknownfor a custom metric - Wiring Prometheus metrics into Kubernetes autoscaling
Not for this skill when
- CPU/memory HPA is enough (simpler, built in)
- Scaling to zero on events (consider KEDA)
- The metric lives outside the cluster with no exporter
Steps
Step 1: Confirm the metrics API the HPA needs is served
kubectl get --raw "/apis/custom.metrics.k8s.io/v1beta1" | head -c 300
kubectl get --raw "/apis/external.metrics.k8s.io/v1beta1" | head -c 300Expected: JSON listing available metrics for each API. If the API itself 404s, no adapter is installed; install Prometheus Adapter (custom metrics) or KEDA (external metrics) first.
Step 2: Verify your metric is actually served
kubectl get --raw "/apis/custom.metrics.k8s.io/v1beta1/namespaces/[namespace]/pods/*/http_requests_per_second" | jq .items[0].valueExpected: a numeric value. If the metric name is absent, the adapter's rules do not map it; fix the adapter's config (metric naming rules, series filters) before touching the HPA.
Step 3: Write the HPA against the verified metric name
apiVersion: autoscaling/v2
kind: HorizontalPodAutoscaler
metadata:
name: [app]-hpa
namespace: [namespace]
spec:
scaleTargetRef:
apiVersion: apps/v1
kind: Deployment
name: [deployment]
minReplicas: 2
maxReplicas: 20
metrics:
- type: Pods
pods:
metric:
name: http_requests_per_second
target:
type: AverageValue
averageValue: "100"Expected: kubectl apply succeeds and kubectl describe hpa shows the metric with a current value, not unknown.
Step 4: Load-test and watch scaling behavior
kubectl get hpa [app]-hpa -n [namespace] -wExpected: TARGETS column shows current/average values and REPLICAS moves within min/max as load changes. If replicas never move, the target value is unreachable (too high/low) or the metric is stale.
Variant phrasings
"hpa custom metric unknown"
The adapter is not serving that metric name. Step 2 is the check; fix adapter rules, not the HPA.
"kubernetes autoscale on queue depth"
Expose queue depth via Prometheus, map it in the adapter, then HPA type Pods or Object as in step 3.
Why it happens
The HPA controller only speaks the metrics APIs; it cannot read Prometheus directly. The adapter is the translator, and its naming rules decide which Prometheus series become which metric names. When names mismatch, the HPA sees unknown and never scales, which looks like an HPA bug but is an adapter config gap.
Edge cases and pitfalls
- Metric staleness: if the series stops (no traffic), the adapter may serve the last value forever; set adapter rules to drop stale series.
- Scaling on a per-pod metric with
type: Podsaverages across pods; a single hot pod will not trigger scale-up the way you expect. - HPA has a default 5-minute scale-down stabilization; custom metrics do not change that, so expect slow scale-down.
- Do not run HPA and VPA auto mode on the same workload's CPU; they fight. Custom-metric HPA plus VPA on memory is the safe pairing.
Provenance
Resolved from the public thread: https://vectle.com/posts/pst_T8zjRBpe43z8tDvhS2crhQ