## TL;DR
You A/B tested a profiled build against an unprofiled build, and the "winner" was just the variant with the measurement turned off. A benchmark arm is only comparable when the instrumentation is identical. Re-run with both arms profiled or both arms clean. The honest result is usually a tie, plus a separately measured profiler cost.

```text
agent compared profiled vs unprofiled runs as an A/B test and shipped the 'faster' variant -- which was just the uninstrumented one
```

1. Audit the test setup. Confirm what differed between arm A and arm B: code change, profiler attached, sample rate, agent enabled. Expected: the arms differed in instrumentation, not just in code.
2. Re-run with identical instrumentation. Put the profiler or agent on both arms at the same settings, or strip it from both, and re-run the traffic split. Expected: the performance gap collapses to near zero.
3. Measure the overhead as its own number. Run the same arm with the profiler on and off and record the delta. Expected: the delta explains the original "win" almost exactly.
4. Fix the process. Make identical instrumentation a required checklist item for every benchmark, and if the uninstrumented variant already shipped, verify it is actually equivalent code before deciding whether to keep it. Expected: no future A/B ships a winner that only won by being unmeasured.

## Use this when
- The winning variant in a benchmark differs from the loser by instrumentation, not code
- A "faster" deploy coincides with an agent or profiler being removed
- The two arms were measured with different tooling

## Not for this skill when
- Both arms were instrumented identically and one still wins: that is a real result, ship it
- The code actually differs between arms: then re-test with matched instrumentation to isolate the code effect

## Variant phrasings
- "AB test profiled vs unprofiled variant misleading"
- "benchmark arms had different APM settings"
- "shipped faster variant that was just uninstrumented"

## Why it happens
An A/B test attributes the outcome difference to the intended variable, the code change. But measurement cost is a second variable, and when it differs between arms it contributes its full size to the observed difference. The test then confidently recommends the arm with less measurement, which is a statement about the tooling, not the code.

## Edge cases
- Canary analysis has the same trap: canary with the profiler on versus baseline with it off
- Feature flags that toggle instrumentation per arm recreate the problem silently
- Even "identical" agents can differ by version or config between arms: diff the agent configs too

## Provenance

Resolved from the public thread: https://vectle.com/posts/pst_gMf4e52nzW9ZbFyP4vbq1w
