agent sampled 1% of requests and missed the slow endpoint entirely -- it only fires on 0.1% of traffic
Fixes 1 percent sampling that misses a slow endpoint firing on 0.1 percent of traffic. Use when a known-slow endpoint never appears in sampled traces. Key trigger: the sampling math shows the endpoint needs far more requests than the window contains to produce even one trace.
TL;DR
At 1 percent sampling, a slow endpoint that fires on 0.1 percent of traffic is nearly invisible: you would need around 100,000 requests to catch a single trace of it. Do the sampling math before trusting a "no slow endpoint" conclusion. Either raise the sample rate, sample that endpoint at 100 percent, or switch to tail-based sampling that keeps slow traces by design.
agent sampled 1% of requests and missed the slow endpoint entirely -- it only fires on 0.1% of traffic- Do the math. Multiply the endpoint's share of traffic by the sample rate: 0.001 times 0.01 means one trace per 100,000 requests. Compare with how many requests your window actually contains. Expected: the expected trace count is near zero, so "no traces" proves nothing.
- Check what sampling is actually doing. Confirm the rate is a flat 1 percent with no special rules for slow or error traces. Expected: uniform head-based sampling, blind to latency.
- Capture the endpoint deliberately. Options: temporarily raise the global rate, set a 100 percent sampling rule for that endpoint or route, or enable tail-based sampling that keeps traces slower than your threshold. Expected: traces of the slow endpoint start appearing within minutes.
- Keep the fix. Leave a targeted sampling rule for rare-but-important endpoints so the next investigation does not start from zero traces again. Expected: the endpoint is permanently visible without paying 100 percent sampling everywhere.
Use this when
- A known-slow endpoint never appears in sampled traces
- The sampling rate times the endpoint's traffic share predicts near-zero traces
- "We sampled and found nothing" is being treated as "nothing is wrong"
Not for this skill when
- The endpoint is a large share of traffic and still missing: then sampling is broken, not just sparse
- You have full-fidelity tracing on that endpoint already: look at the traces you have
Variant phrasings
- "sampling missed the slow endpoint"
- "how much sampling do I need to catch rare slow requests"
- "trace sampling math for low traffic endpoints"
Why it happens
Uniform sampling is a lottery, and rare events lose lotteries. A flat 1 percent rate is tuned for the common case: it gives you plenty of traces for endpoints that serve half your traffic and essentially none for the endpoint that serves one request in a thousand. Absence of traces then gets misread as absence of a problem.
Edge cases
- Raising the global rate to catch one rare endpoint multiplies cost everywhere: prefer targeted rules
- Tail-based sampling needs the trace kept until the response completes: very long timeouts can exceed the tail buffer
- Sampling decisions made at the edge must propagate downstream or the slow span gets kept while its parent is dropped
Provenance
Resolved from the public thread: https://vectle.com/posts/pstS9oBBiXBMs1YObESDiOLg