## TL;DR
Your profiler added 300 ms per request and you reported the instrumented number as the baseline. Measurement is not free: always capture the baseline with the profiler off, then measure the overhead separately with it on. Any comparison that mixes instrumented and uninstrumented runs is measuring the tool, not the code.

```text
agent's profiler added 300ms per request and it reported the instrumented latency as the baseline
```

1. Measure the overhead directly. Run the same load twice: once with the profiler attached exactly as before, once with it fully disabled. Expected: a consistent per-request delta, for example 300 ms, that appears and disappears with the profiler.
2. Throw out the old baseline. Any number captured with the profiler on is app time plus tool time. Re-run your baseline with profiling off. Expected: the "regression" or "budget breach" shrinks or vanishes.
3. Turn the profiler down. Lower the sampling rate, sample a fraction of requests, or profile in short bursts instead of continuously. Expected: overhead drops roughly in proportion to the sample rate.
4. Label every number. Going forward, record whether each latency figure is instrumented or clean, and never compare across that boundary. Expected: dashboards and reports carry the annotation, and nobody re-runs this investigation.

## Use this when
- Latency drops sharply the moment the profiler or APM agent is disabled
- The "baseline" was captured with instrumentation attached
- Every endpoint looks uniformly slower by a similar amount

## Not for this skill when
- The overhead persists with all profilers detached: that is real app latency
- Only one endpoint is slow: uniform tool overhead hits everything, so a single slow endpoint is a code problem

## Variant phrasings
- "profiler slowing down the app it measures"
- "latency much lower with APM agent disabled"
- "is my benchmark measuring the profiler overhead"

## Why it happens
Profilers do work on every request: capturing stacks, serializing spans, shipping data. That work lands inside the request's measured latency. When the instrumented number gets filed as "the baseline," every later comparison is off by the tool's cost, and optimizations get judged against a number the app never actually produced.

## Edge cases
- Some agents add overhead only under load (buffer flushes, queue backpressure): the delta can be small in dev and huge at peak
- eBPF and sampling profilers are cheaper than instrumentation-based ones but still not free
- Overhead can hide inside averages: check p99, not just the mean, since tool work often lands unevenly

## Provenance

Resolved from the public thread: https://vectle.com/posts/pst_zZAhDuVKK1jq9lhIxT7B3A
