how to debug Prometheus target discovery
Debugs Prometheus target discovery when targets show down or missing. Use when Prometheus has no targets, targets flap between up and down, or new services never appear as targets. Covers service discovery configs, relabeling, and scrape errors. Not for PromQL query issues, alert rule problems, or remote-write setup.
TL;DR
When Prometheus targets are missing or down, the problem is in one of three places: discovery (Prometheus never learns about the target), relabeling (it learns but drops or mangles it), or the scrape itself (it tries and fails). Check the Targets page in that order: discovery config, then dropped targets, then the scrape error message. The error column on the Targets page usually names the fix.
Error / query
how to debug Prometheus target discoveryUse this skill when
- Prometheus shows no targets or targets down
- New services never appear as scrape targets
- Targets flap between up and down
- The Targets page shows errors you need to decode
Not for this skill when
- Queries return no data but targets are up (query problem)
- Alert rules misbehave (rule problem)
- Setting up remote write (different config)
Steps
Step 1: Check the discovery configuration
kubectl get prometheus [prom-name] -n [monitoring] -o yaml | grep -A15 "serviceMonitorSelector\|podMonitorSelector"
# or for static config: check scrape_configs job definitionsExpected: the selectors Prometheus uses to find targets. If a ServiceMonitor's labels do not match the selector, the targets never enter discovery; this is the most common no targets cause.
Step 2: Look at dropped targets
Prometheus UI: Status, then Targets. Scroll to the dropped-targets section
(or query up==0 and check the discovery labels).Expected: targets that discovery found but relabeling dropped. Dropped targets mean the relabel config is filtering them; fix the keep/drop rules rather than the discovery.
Step 3: Read the scrape error on down targets
On the Targets page, the Error column shows the last scrape failure:
"connection refused" (target not listening), "context deadline exceeded"
(scrape timeout), "404" (wrong metrics path), certificate errors (TLS).Expected: each error maps to a fix: wrong port/path in the ServiceMonitor, scrape timeout too short for a slow endpoint, or TLS config mismatch. Fix the target definition, not Prometheus.
Step 4: Verify end to end after the fix
# force a config reload if needed, then:
# Prometheus UI Targets page: target shows UP with a recent last-scrape timeExpected: the target flips to UP and up{job="[job]"} == 1 in a query. If it flaps, the target itself is unstable (restarting pods, slow endpoint); discovery is fine.
Variant phrasings
"prometheus no targets"
Steps 1-2. Selector mismatch or relabel drops cover nearly all cases.
"prometheus target down connection refused"
Step 3. The target is not listening where the ServiceMonitor says it is.
Why it happens
Prometheus target pipeline has three stages that fail independently: service discovery finds candidates, relabeling filters and rewrites them, and scraping fetches metrics. Each stage's failure looks like no data, but the fix lives at the specific stage. The Targets page exposes all three, which is why it is always step one.
Edge cases and pitfalls
- Pod restarts change IPs; if targets flap on deploys, the scrape interval may just be catching the restart window. Check pod stability first.
- Very large target counts slow discovery; the UI may lag behind reality by a scrape interval or two.
- TLS-enabled endpoints need the right scheme and CA in the scrape config;
httpvshttpsmismatch is a classic silent failure. - Federation and remote-write do not create targets; if you expect federated data, check the federation job, not discovery.
Provenance
Resolved from the public thread: https://vectle.com/posts/pst_oqjv9m4o0milBLvQ1CUMkQ
Maintainer review
No maintainer verification is recorded for this version.
This records the version a maintainer checked. It does not assert that the version is the latest upstream release.