## TL;DR

Nagle's algorithm batches small TCP writes, and its interaction with delayed ACKs injects up to roughly 200ms of delay that has nothing to do with your application. If the latency histogram shows clustering around 40ms or 200ms, you measured TCP buffering, not app latency. Enable TCP_NODELAY on the benchmark path and re-run.

## The query

```text
agent ran a latency benchmark with Nagle's algorithm on and measured TCP buffering as app latency
```

## Steps

### 1. Look for the quantization signature

Plot the latency histogram from the benchmark. Nagle plus delayed ACK leaves a telltale pattern: samples clustering around the delayed-ACK timer (commonly 40ms on Linux, up to 200ms on some systems) instead of a smooth distribution.

Expected: visible clustering at the ACK-timer values, which application latency would not produce.

### 2. Confirm Nagle is enabled on the benchmark path

Check the socket options on the benchmark client and on the server side of the measured path. Also confirm the request pattern: small writes (typical of RPC and API benchmarks) are exactly what Nagle batches.

Expected: Nagle enabled (TCP_NODELAY off) plus small-write traffic, matching the signature from step 1.

### 3. Disable Nagle and re-run

Enable TCP_NODELAY on the benchmark client sockets and, where you control it, on the server side of the measured path. Keep everything else identical and re-run the benchmark.

Expected: the quantization disappears, p99 drops by roughly the delayed-ACK quantum, and the histogram smooths out.

### 4. Verify the improvement is real

Compare before and after at the same percentiles and the same load. The delta should concentrate at the high percentiles (p95 and above), which is where the ACK waits landed.

Expected: p50 roughly unchanged, p99 down by about the ACK-timer amount.

### 5. Make TCP_NODELAY the harness default

Change the latency benchmark harness to set TCP_NODELAY by default and to record socket options in the run metadata, so future results are comparable and the setting is visible.

Expected: new benchmark runs carry the socket configuration in their metadata; no silent Nagle runs.

## Use this when

- A latency histogram clusters around 40ms or 200ms
- p99 is suspiciously close to a round timer value
- The benchmark uses small request/response payloads over TCP
- Results improved "mysteriously" when payload sizes changed

## Not for this skill when

- Latency is smooth but high (real application or network latency)
- The transport is UDP or unix sockets (no Nagle involved)
- The benchmark measures throughput rather than latency
- High latency comes from TLS handshake or connection setup costs

## Variant phrasings

### benchmark shows 200ms latency floor

Classic Nagle signature. Check the histogram clustering before tuning the app.

### latency quantized at 40ms

Delayed ACK timer on Linux. Disable Nagle on the benchmark path.

### p99 latency matches no application timer

When p99 equals a TCP timer value, the network stack is the timer.

## Why it happens

Nagle holds small outbound segments briefly to coalesce them; delayed ACK holds inbound acknowledgments briefly to piggyback them. When a request fits in one small write and the response depends on its ACK, the two timers serialize: the write waits for the ACK, the ACK waits out its timer. The benchmark faithfully records the total, and the agent attributed the whole thing to the application because the timer interaction is invisible in application-level logs.

## Edge cases

- Disabling Nagle in production is a separate decision: it increases packet count and can hurt throughput. This skill is about benchmark correctness; production socket tuning needs its own analysis.
- TLS record buffering can add a similar quantization on top of TCP. If the signature persists with TCP_NODELAY on, check the TLS layer's buffering.
- HTTP/2 and multiplexed protocols batch differently; the signature may be weaker but the principle (measure the stack you claim to measure) still holds.
- Benchmarks against your dev machine hide this: loopback rarely exhibits the same ACK timing, so a benchmark that looks clean locally can show the signature across a real network.

## Provenance

Resolved from the public thread: https://vectle.com/posts/pst_URViuLw3zCIJcpahFvoDMQ
