Jaeger "span not found" for recent traces: retention vs sampling
Diagnoses Jaeger span-not-found errors for recent traces. Use when trace lookup fails for spans that should exist, when deciding between retention and sampling causes, or when trace IDs come from logs. Not for instrumentation issues.
TL;DR
"Span not found" for a recent trace means the span never reached Jaeger or was already dropped: check sampling first (was this trace sampled at all), then retention (is the backend already expiring it), then the trace ID itself (log-propagated IDs are often malformed). Recent plus not-found is usually sampling, not retention.
The query
Jaeger "span not found" for recent traces: retention vs samplingUse this when
- Trace lookup fails for recent trace IDs
- Distinguishing sampling drops from retention expiry
- Trace IDs copied from logs do not resolve
- After changing sampling configuration
Not for when
- Instrumentation not producing spans at all
- Old traces past retention (expected)
- Collector pipeline issues (different stage)
Steps
Step 1: Validate the trace ID format
Check the trace ID for correct length and hex format. IDs copied from logs get truncated or mangled; an invalid ID never matches anything. This is the cheapest check and a common cause. Expected output: the ID confirmed well-formed, or the mangling spotted.
Step 2: Check whether the trace was sampled
Verify the sampling decision for the request: check the sampler configuration and whether this trace type is sampled. Unsampled traces leave no spans; "not found" is the correct answer for them. Expected output: the sampling decision known; unsampled traces explained.
Step 3: Check retention settings
Confirm the backend's retention window covers the trace's age. Aggressive retention (hours, not days) makes even recent traces vanish. Match the retention to how far back people actually look up traces. Expected output: retention window known and adequate, or identified as the cause.
Step 4: Verify the write path for that service
Check that the service's spans reach the backend: collector logs, exporter metrics for that service. A service whose exporter is broken produces "not found" for all its traces, which looks like a lookup problem. Expected output: the service's export path confirmed working.
Step 5: Align sampling with lookup expectations
If people routinely look up traces that sampling drops, the sampling policy is wrong for the use case: keep all error traces, sample the rest. Sampling policy should serve the debugging workflow, not just the budget. Expected output: the traces people need are the traces that get kept.
Provenance
Resolved from the public thread: https://vectle.com/posts/pst_Fob5peYpIR-BPbuINmi-0g
Maintainer review
No maintainer verification is recorded for this version.
This records the version a maintainer checked. It does not assert that the version is the latest upstream release.