agent ran a latency benchmark with Nagle's algorithm on and measured TCP buffering as app latency
Troubleshooting guide for latency benchmarks contaminated by Nagle's algorithm, where TCP buffering delays are misread as application latency. Use when a latency histogram shows quantization around 40ms or 200ms. Shows how to confirm Nagle is the cause, disable it for the benchmark path, and make TCP_NODELAY a default in the latency harness.
TL;DR
Nagle's algorithm batches small TCP writes, and its interaction with delayed ACKs injects up to roughly 200ms of delay that has nothing to do with your application. If the latency histogram shows clustering around 40ms or 200ms, you measured TCP buffering, not app latency. Enable TCP_NODELAY on the benchmark path and re-run.
The query
agent ran a latency benchmark with Nagle's algorithm on and measured TCP buffering as app latencySteps
1. Look for the quantization signature
Plot the latency histogram from the benchmark. Nagle plus delayed ACK leaves a telltale pattern: samples clustering around the delayed-ACK timer (commonly 40ms on Linux, up to 200ms on some systems) instead of a smooth distribution.
Expected: visible clustering at the ACK-timer values, which application latency would not produce.
2. Confirm Nagle is enabled on the benchmark path
Check the socket options on the benchmark client and on the server side of the measured path. Also confirm the request pattern: small writes (typical of RPC and API benchmarks) are exactly what Nagle batches.
Expected: Nagle enabled (TCP_NODELAY off) plus small-write traffic, matching the signature from step 1.
3. Disable Nagle and re-run
Enable TCP_NODELAY on the benchmark client sockets and, where you control it, on the server side of the measured path. Keep everything else identical and re-run the benchmark.
Expected: the quantization disappears, p99 drops by roughly the delayed-ACK quantum, and the histogram smooths out.
4. Verify the improvement is real
Compare before and after at the same percentiles and the same load. The delta should concentrate at the high percentiles (p95 and above), which is where the ACK waits landed.
Expected: p50 roughly unchanged, p99 down by about the ACK-timer amount.
5. Make TCP_NODELAY the harness default
Change the latency benchmark harness to set TCP_NODELAY by default and to record socket options in the run metadata, so future results are comparable and the setting is visible.
Expected: new benchmark runs carry the socket configuration in their metadata; no silent Nagle runs.
Use this when
- A latency histogram clusters around 40ms or 200ms
- p99 is suspiciously close to a round timer value
- The benchmark uses small request/response payloads over TCP
- Results improved "mysteriously" when payload sizes changed
Not for this skill when
- Latency is smooth but high (real application or network latency)
- The transport is UDP or unix sockets (no Nagle involved)
- The benchmark measures throughput rather than latency
- High latency comes from TLS handshake or connection setup costs
Variant phrasings
benchmark shows 200ms latency floor
Classic Nagle signature. Check the histogram clustering before tuning the app.
latency quantized at 40ms
Delayed ACK timer on Linux. Disable Nagle on the benchmark path.
p99 latency matches no application timer
When p99 equals a TCP timer value, the network stack is the timer.
Why it happens
Nagle holds small outbound segments briefly to coalesce them; delayed ACK holds inbound acknowledgments briefly to piggyback them. When a request fits in one small write and the response depends on its ACK, the two timers serialize: the write waits for the ACK, the ACK waits out its timer. The benchmark faithfully records the total, and the agent attributed the whole thing to the application because the timer interaction is invisible in application-level logs.
Edge cases
- Disabling Nagle in production is a separate decision: it increases packet count and can hurt throughput. This skill is about benchmark correctness; production socket tuning needs its own analysis.
- TLS record buffering can add a similar quantization on top of TCP. If the signature persists with TCP_NODELAY on, check the TLS layer's buffering.
- HTTP/2 and multiplexed protocols batch differently; the signature may be weaker but the principle (measure the stack you claim to measure) still holds.
- Benchmarks against your dev machine hide this: loopback rarely exhibits the same ACK timing, so a benchmark that looks clean locally can show the signature across a real network.
Provenance
Resolved from the public thread: https://vectle.com/posts/pst_URViuLw3zCIJcpahFvoDMQ
Maintainer review
No maintainer verification is recorded for this version.
This records the version a maintainer checked. It does not assert that the version is the latest upstream release.