nginx "upstream prematurely closed connection": 502 debugging
Fixes nginx 502s from upstreams closing connections early. Use when nginx logs upstream prematurely closed connection, when backends die mid-request, or when keepalive mismatches occur. Not for 502s from DNS or refused connections.
TL;DR
The upstream accepted the connection but closed it before responding: usually the backend crashed, timed out internally, or has keepalive settings incompatible with nginx's (the backend closes idle keepalive connections exactly when nginx reuses them). Check the backend's logs for the crash, then align keepalive timeouts so nginx never reuses a connection the backend is about to close.
The query
nginx "upstream prematurely closed connection": 502 debuggingUse this when
- nginx error log shows upstream prematurely closed connection
- 502s correlate with backend deploys or restarts
- Keepalive-related 502s under load
- Backend seems healthy but nginx 502s
Not for when
- 502 from connection refused (backend down)
- DNS resolution failures for upstreams
- SSL errors to upstreams
Steps
Step 1: Correlate with backend logs
Match the 502 timestamps against the backend's logs. Crashes, panics, or OOM kills at the same time explain the closed connections directly; the fix is the backend bug, not nginx tuning. Expected output: backend crashes confirmed or ruled out.
Step 2: Check keepalive timeout alignment
Compare nginx's keepalive timeout with the backend's keepalive/idle timeout. If the backend closes idle connections sooner than nginx expects, nginx reuses a dead connection and gets exactly this error. The backend's timeout must exceed nginx's. Expected output: timeouts aligned with the backend's comfortably above nginx's.
Step 3: Verify during backend restarts
Deploys that restart backends produce bursts of these 502s as connections drain. Confirm the errors cluster around deploys; if so, the fix is graceful shutdown (drain timeouts) on the backend, not nginx config. Expected output: deploy-correlated 502s addressed with graceful shutdown.
Step 4: Test with keepalive disabled temporarily
Disable upstream keepalive briefly to see if the 502s stop. If they do, the issue is definitively keepalive lifecycle, not request handling. Re-enable with corrected timeouts. Expected output: the keepalive hypothesis proven or killed.
Step 5: Monitor 502 rate as the deploy signal
Track the 502 rate continuously; it should be near zero outside deploys. A rising baseline means the backend is degrading, and the 502s are the early warning, not the problem to tune away. Expected output: 502s as a monitored signal, not background noise.
Provenance
Resolved from the public thread: https://vectle.com/posts/pst_8UnGItZZ9OkgRWLX04ROcQ