## TL;DR
Correlation is just a shared ID: the trace ID should appear in every log line, and the same timestamp window should appear in your metrics query. Start from the alert, grab the trace ID from the log, open the trace, then pull metrics for the same 15-minute window. If your logs do not carry trace IDs, fix that first; nothing else works without it.

## Error / query
```text
how to correlate logs traces and metrics in an incident
```

## Use this skill when
- An alert fired and you are bouncing between three tools trying to connect the dots
- Logs show errors but you cannot tell which requests or which downstream call caused them
- You are setting up instrumentation and want the three signals to link from day one
- A postmortem needs a single timeline across logs, traces, and metrics

## Not for this skill when
- The incident is a single host problem (disk full, OOM); you need host metrics, not correlation
- Your services do not emit trace IDs yet; add instrumentation first, then come back
- You are doing long-term trend analysis; correlation is an incident-time technique

## Steps

### Step 1: Pull the trace ID out of the error logs
```bash
journalctl -u myservice --since "30 minutes ago" -p err | grep -o 'trace_id=[a-f0-9]*' | sort | uniq -c | sort -rn | head -5
```
Expected: one or two trace IDs dominating the errors; the top one is your incident thread to pull.

### Step 2: Open that trace and find the slow or failing span
```bash
TRACE=$(journalctl -u myservice --since "30 minutes ago" -p err | grep -o 'trace_id=[a-f0-9]*' | head -1 | cut -d= -f2); echo "trace: $TRACE"
```
Expected: a single trace ID printed; paste it into your trace viewer to see the full request path and which span failed.

### Step 3: Pull metrics for the same time window as the trace
```bash
curl -s -G 'https://example.com/api/v1/query_range' --data-urlencode 'query=sum(rate(http_request_duration_seconds_count{service="myservice"}[5m]))' --data-urlencode 'start=2026-10-04T04:00:00Z' --data-urlencode 'end=2026-10-04T04:15:00Z' --data-urlencode 'step=60s' | head -c 500
```
Expected: a request-rate series for the incident window; look for the dip or spike matching the trace timestamps.

### Step 4: Check downstream dependency metrics in that window
```bash
curl -s -G 'https://example.com/api/v1/query' --data-urlencode 'query=sum(rate(http_requests_total{service="myservice",status=~"5.."}[15m])) by (downstream)' | head -c 500
```
Expected: error counts broken down by downstream; the failing dependency stands out.

### Step 5: Confirm the log volume matches the metric signal
```bash
journalctl -u myservice --since "30 minutes ago" -p err | wc -l
```
Expected: the error log count roughly matches the metric error count; a big mismatch means logs or metrics are dropping data.

## Variant phrasings

### "How to go from a metric alert to the bad trace"
Take the alert's service and time window, query logs for errors in that window, extract the most common trace ID, open it. Metric to log to trace, in that order.

### "Correlating logs and traces without a trace ID"
Use timestamp plus a request-scoped field you do have (user ID, order ID); it is slower and fuzzier, and it is the reason to add trace IDs.

### "Exemplars in Prometheus"
Exemplars attach a trace ID to a metric sample, so the metric itself links to the trace; if your setup supports them, they skip step 1 entirely.

## Why it happens
Logs, traces, and metrics are usually built by different people at different times, so they share nothing by default. The trace ID is the cheapest join key: one field, added at request ingress, propagated everywhere, and every signal becomes one click away from the others.

## Edge cases and pitfalls
- Clock skew between hosts makes timestamp correlation lie; keep NTP healthy or the windows will not line up.
- Async and batch jobs have no single request trace; correlate on job ID instead.
- High-cardinality trace IDs in metric labels will explode your metrics store; keep the ID in exemplars or logs, never as a label.
- Sampled-out traces break the chain; during an incident, raise sampling for the affected service.

## Provenance

Resolved from the public thread: https://vectle.com/posts/pst_C0o41dIzgQYOeN6S9p4tww
