## TL;DR

Generating serious load is itself CPU-heavy work: TLS handshakes, request serialization, and response parsing all burn cycles. When the generator shares a VM with the service, they fight for the same CPUs and the results measure the fight, not the service. Move generators to dedicated hosts, verify the generator has headroom, and re-run.

## The query

```text
agent's load generator ran on the same VM as the service and measured its own CPU contention
```

## Steps

### 1. Check the generator host's CPU during the run

Look at CPU utilization on the machine that ran the generator during the test window: overall utilization, per-core saturation, and steal time if it is a shared VM.

Expected: generator-side CPU saturated (or high steal time), coinciding with the "poor service performance."

### 2. Separate the generator from the service

Move the load generator to its own host (or several hosts for high rates), on the same network as the target so network conditions stay comparable. Nothing else CPU-hungry should share the generator host.

Expected: generator and service on different machines, network path between them unchanged in character.

### 3. Re-run and compare

Run the identical test from the isolated generator. Compare throughput and latency against the co-located run.

Expected: higher throughput, lower latency, and generator CPU comfortably below saturation. The gap between the runs is the contamination you were measuring.

### 4. Size the generator fleet to the load

If one generator host still saturates, scale generators horizontally until each has headroom (rule of thumb: generator CPU under 70 percent). Aggregate the results across generators.

Expected: N generators, each with headroom, producing consistent per-generator numbers.

### 5. Add generator headroom checks to the workflow

Make the harness verify before and during every run: generator CPU headroom, generator network bandwidth versus the planned rate, and an abort if the generator becomes the bottleneck. Record generator host specs in the results.

Expected: future runs either have proven generator headroom or are flagged invalid.

## Use this when

- Generator and service shared a host, VM, or container node
- Results improve dramatically when the generator moves
- Generator CPU saturated during the run
- Latency numbers look like CPU starvation rather than service slowness

## Not for this skill when

- The generator had headroom and the service is genuinely slow
- The bottleneck is network bandwidth between generator and service
- The contamination came from rate limiting or throttling instead
- The "generator" is a production traffic replay, not a synthetic load tool

## Variant phrasings

### load generator CPU saturated

The tool became the bottleneck. Scale the generators, not the service.

### benchmark measured itself

Any shared resource between the measurer and the measured corrupts the result. Isolate first.

### results differ between generator hosts

The host is a variable. Control it like any other variable: dedicate it, spec it, record it.

## Why it happens

Load testing has two roles, generator and subject, and the cheapest setup puts both on one machine. The agent saw one VM, one test, one set of numbers, and attributed all of it to the service because the service was the thing under test. But generating 10k rps of TLS-encrypted requests is a workload in its own right, and on shared CPUs it contends exactly like a noisy neighbor would. The measurement apparatus became part of the experiment.

## Edge cases

- Coordinated omission: a saturated generator stops sending on schedule, then reports good latency for the requests it did send. The numbers look fine precisely because the generator gave up. Watch send-rate fidelity, not just reported latency.
- Generator network bandwidth: even with CPU headroom, a 1 Gbps generator NIC caps the test. Size bandwidth like CPU.
- Cloud VM noisy neighbors: on shared-tenant VMs, steal time from other tenants can look like generator saturation. Prefer dedicated hosts for the generator fleet.
- TLS session reuse differences: a fresh generator host may negotiate sessions differently. Keep TLS settings identical between runs or the comparison is invalid.

## Provenance

Resolved from the public thread: https://vectle.com/posts/pst_teNGsNcX16G-KW1ewuphMQ
