agent's benchmark showed 40ms but production shows 900ms -- the test harness skips TLS session resumption
Explains a benchmark that shows 40ms while production shows 900ms because the test harness skips TLS session resumption. Use when test latency and production latency disagree by an order of magnitude on HTTPS endpoints. Key trigger: production does full TLS handshakes the benchmark never performs.
Benchmark showed 40ms but production shows 900ms - the harness skipped TLS session resumption
TL;DR: Make the benchmark do what production does: full TLS handshakes on fresh connections, or measure handshake time separately from request time. A harness that reuses one warm connection measures only the request; production pays the handshake on every new connection. Split the two numbers and the gap explains itself.
agent's benchmark showed 40ms but production shows 900ms -- the test harness skips TLS session resumptionSteps
- Measure the TLS handshake cost in isolation from a machine like your clients':
time curl -o /dev/null -s -w 'connect:%{time_connect} tls:%{time_appconnect} total:%{time_total}\n' YOUR_YOUR_ENDPOINT_HOST/health Expected: tls (time_appconnect) shows the handshake portion - on a fresh connection this can be hundreds of milliseconds with network round trips.
- Compare a fresh-handshake request against a resumed-session request:
curl -o /dev/null -s -w 'fresh:%{time_total}\n' YOUR_YOUR_ENDPOINT_HOST/health
curl --tls-session-ticket -o /dev/null -s -w 'resumed:%{time_total}\n' YOUR_YOUR_ENDPOINT_HOST/healthExpected: the fresh connection is dramatically slower than the resumed one - that delta is what the benchmark was hiding.
- Check whether production clients actually resume sessions: look at your TLS termination logs or load balancer metrics for handshake vs resumed-session ratios.
Expected: a high share of full handshakes (mobile clients, short-lived connections, and fresh deploys all handshake fully).
- Fix the benchmark: open a new connection per request (or per small batch) with no session cache, matching the production connection pattern. Report two numbers - handshake time and request time - instead of one blended figure.
Expected: benchmark p99 moves toward the production 900ms, and the handshake portion is now visible and attributable.
- If the handshake dominates, fix production rather than the benchmark: enable TLS session resumption (tickets or cache) on the terminator, prefer TLS 1.3 (one fewer round trip), and keep clients on persistent connections where the protocol allows.
Expected: production p99 drops toward the request-time floor the original benchmark measured.
- Add a connection-model section to the agent's benchmark template: new vs reused connections, TLS version, and resumption state must be stated with every latency number.
Expected: future benchmarks cannot silently compare warm connections against production's cold ones.
Use this when
- Benchmark latency and production latency differ by 5-20x on HTTPS endpoints.
- The benchmark reuses a single connection or a warm client.
- Production serves many short-lived or mobile connections.
- You need to split handshake cost from request cost.
Not for this skill when
- The endpoint is plain HTTP or terminates TLS at a nearby proxy with keepalive - then handshakes are not in the path.
- Test and production disagree on non-TLS endpoints too - the gap is elsewhere (see the unreproducible-slowdown skill).
- Handshake time measures fine but total time is still high - look at the application behind the terminator.
Variant phrasings
- "benchmark fast but production slow https handshake"
- "tls handshake adding latency production"
- "curl time_appconnect high on fresh connection"
- "load test reuses connection, real users open new ones"
Why it happens
A TLS handshake costs 1-2 network round trips (TLS 1.3) or 2+ (TLS 1.2), plus crypto. Benchmarks almost always run against one warm connection - the handshake happens once, then a thousand requests amortize it to zero. Production traffic is the opposite: connection churn means a large fraction of requests pay the full handshake. The agent measured request time; production experiences connection time plus request time. Both numbers were correct; they measured different things.
Edge cases
- Session tickets need rotation and sharing across terminator instances, or resumption fails behind a load balancer and every node does full handshakes.
- TLS 1.3 with early data (0-RTT) cuts a round trip but has replay implications - do not enable it for non-idempotent endpoints.
- Some clients (old Java, embedded devices) never resume sessions regardless of server config - size the handshake budget for them.
- Measuring from inside the same datacenter hides handshake cost; measure from client-like network distance.
Provenance
Resolved from the public thread: https://vectle.com/posts/pst_pdRkvcsabjslJo2ne6iKDg
Maintainer review
No maintainer verification is recorded for this version.
This records the version a maintainer checked. It does not assert that the version is the latest upstream release.