## TL;DR
When Prometheus targets are missing or down, the problem is in one of three places: discovery (Prometheus never learns about the target), relabeling (it learns but drops or mangles it), or the scrape itself (it tries and fails). Check the Targets page in that order: discovery config, then dropped targets, then the scrape error message. The error column on the Targets page usually names the fix.

## Error / query
```text
how to debug Prometheus target discovery
```

## Use this skill when
- Prometheus shows no targets or targets down
- New services never appear as scrape targets
- Targets flap between up and down
- The Targets page shows errors you need to decode

## Not for this skill when
- Queries return no data but targets are up (query problem)
- Alert rules misbehave (rule problem)
- Setting up remote write (different config)

## Steps

### Step 1: Check the discovery configuration
```bash
kubectl get prometheus [prom-name] -n [monitoring] -o yaml | grep -A15 "serviceMonitorSelector\|podMonitorSelector"
# or for static config: check scrape_configs job definitions
```
Expected: the selectors Prometheus uses to find targets. If a ServiceMonitor's labels do not match the selector, the targets never enter discovery; this is the most common `no targets` cause.

### Step 2: Look at dropped targets
```text
Prometheus UI: Status, then Targets. Scroll to the dropped-targets section
(or query up==0 and check the discovery labels).
```
Expected: targets that discovery found but relabeling dropped. Dropped targets mean the relabel config is filtering them; fix the keep/drop rules rather than the discovery.

### Step 3: Read the scrape error on down targets
```text
On the Targets page, the Error column shows the last scrape failure:
"connection refused" (target not listening), "context deadline exceeded"
(scrape timeout), "404" (wrong metrics path), certificate errors (TLS).
```
Expected: each error maps to a fix: wrong port/path in the ServiceMonitor, scrape timeout too short for a slow endpoint, or TLS config mismatch. Fix the target definition, not Prometheus.

### Step 4: Verify end to end after the fix
```bash
# force a config reload if needed, then:
# Prometheus UI Targets page: target shows UP with a recent last-scrape time
```
Expected: the target flips to UP and `up{job="[job]"} == 1` in a query. If it flaps, the target itself is unstable (restarting pods, slow endpoint); discovery is fine.

## Variant phrasings

### "prometheus no targets"
Steps 1-2. Selector mismatch or relabel drops cover nearly all cases.

### "prometheus target down connection refused"
Step 3. The target is not listening where the ServiceMonitor says it is.

## Why it happens
Prometheus target pipeline has three stages that fail independently: service discovery finds candidates, relabeling filters and rewrites them, and scraping fetches metrics. Each stage's failure looks like `no data`, but the fix lives at the specific stage. The Targets page exposes all three, which is why it is always step one.

## Edge cases and pitfalls
- Pod restarts change IPs; if targets flap on deploys, the scrape interval may just be catching the restart window. Check pod stability first.
- Very large target counts slow discovery; the UI may lag behind reality by a scrape interval or two.
- TLS-enabled endpoints need the right scheme and CA in the scrape config; `http` vs `https` mismatch is a classic silent failure.
- Federation and remote-write do not create targets; if you expect federated data, check the federation job, not discovery.

## Provenance

Resolved from the public thread: https://vectle.com/posts/pst_oqjv9m4o0milBLvQ1CUMkQ
