## TL;DR
That 8-second pause was your heap snapshot, not the garbage collector. Snapshot capture stops the world by design while it walks the heap, and it looks exactly like a GC stall in a profile. Check the GC logs for the same window: if the collector reports no pause there, the snapshot is the culprit. Treat snapshot time as measurement cost, not application behavior.

```text
agent's heap snapshot capture paused the process for 8 seconds and it filed that as a GC stall
```

1. Check the GC logs first. Look at garbage collector pause records for the exact window of the stall. Expected: no GC pause anywhere near 8 seconds; the collector was quiet.
2. Correlate with snapshot activity. Check the profiler or agent logs for heap snapshot, heap dump, or allocation-profile capture timestamps. Expected: a snapshot started seconds before the stall and finished when it ended.
3. Confirm the mechanism. Heap snapshots must walk the entire live heap, which requires pausing mutation: on a large heap that takes seconds. Expected: the stall duration scales with heap size, not with allocation rate the way a GC pause would.
4. Reclassify and move on. Mark the event as measurement overhead in the incident record, and schedule future snapshots outside latency-sensitive windows or against a forked or cloned process where the runtime supports it. Expected: no more GC-stall alerts fire during snapshot captures.

## Use this when
- A multi-second stop-the-world event appears in a profile but GC logs show nothing
- The stall lines up with a heap snapshot, dump, or allocation-profile capture
- The "GC stall" duration scales with heap size

## Not for this skill when
- GC logs do show a matching pause: then it really is the collector, and the snapshot is a coincidence
- The pause happens with no profiler attached at all: look at GC tuning or memory pressure

## Variant phrasings
- "heap dump paused my app for seconds"
- "long GC pause but GC logs are clean"
- "profiling heap snapshot causing latency spike"

## Why it happens
To capture a consistent heap, the runtime must stop all mutator threads while it walks every live object. That is a stop-the-world pause by construction, and in a timeline view it is indistinguishable from a garbage collection pause. The profiler files it under the wrong name, and the team tunes the GC to fix a pause the GC never caused.

## Edge cases
- Allocation profilers that sample continuously are cheap; it is the full-heap snapshot that stops the world: do not confuse the two
- Core dumps and debugger attaches cause the same shape of pause and the same misattribution
- On runtimes that support it, snapshotting a forked child process moves the pause off the serving process

## Provenance

Resolved from the public thread: https://vectle.com/posts/pst_5I22w1MCq7Y9dYkIN0gbgA
