agent's load test DDoS'd the staging database -- it ramped to 10k rps without checking the connection pool size first
Recovery and prevention guide for load tests that exhaust the database connection pool by ramping to high request rates without checking pool capacity first. Use after a staging database was overwhelmed by a load test. Shows how to size required concurrency, compare against pool limits, and re-run with a guarded ramp that watches pool waiters.
TL;DR
10k requests per second needs roughly rps times average latency in concurrent connections, and most databases have a few hundred slots. The test opened connections the database could never serve. Size the concurrency first, compare it against the pool and database limits, add a pooler if needed, and re-run with a gradual ramp that aborts when pool waiters appear.
The query
agent's load test DDoS'd the staging database -- it ramped to 10k rps without checking the connection pool size firstSteps
1. Stop the test and assess the database
Halt the load run immediately. Check the database for connection exhaustion: current connection count versus the max connections setting, and any "remaining connection slots are reserved" errors in the log.
Expected: confirmation of whether the database hit its connection ceiling, and that it recovers once the test stops.
2. Compute the concurrency the target rate requires
Required concurrent connections is roughly target rps multiplied by average query latency in seconds. 10k rps at 50ms average latency needs about 500 concurrent connections, before any headroom.
Expected: a concrete concurrency number for the planned load.
3. Compare against the pool and the database
Read the actual limits: the database max connections setting, the application pool size, and the pooler configuration if one sits in front (pool size, pool mode). If required concurrency exceeds any of them, the test as planned cannot run.
Expected: a clear statement of which limit the test would have breached.
4. Close the gap before re-running
Options: lower the target rps to fit the pool, raise the pool within safe bounds, or put a transaction-mode pooler in front of the database so thousands of client connections multiplex onto hundreds of server connections. Pick the combination that fits staging's purpose.
Expected: required concurrency fits comfortably inside every limit, with headroom.
5. Re-run with a guarded ramp
Ramp in stages (for example 10 percent, 25 percent, 50 percent, 100 percent of target), holding each stage while watching pool waiter metrics: queued checkouts, wait events on the database, and connection errors. Abort the run if waiters grow.
Expected: the full target rate sustains with pool waiters near zero and no connection errors.
Use this when
- A load test caused connection exhaustion or "too many clients" errors
- The test ramped straight to a high rps target with no pre-flight
- Pool waiters or connection timeouts appeared during the run
- The agent planned the test from rps alone without a concurrency budget
Not for this skill when
- The database is slow for query-plan reasons rather than connection limits
- The pool is sized fine and the bottleneck is CPU or disk
- The test targeted shared production infrastructure (isolate the target instead)
- Failures are 429s from a rate limiter rather than connection errors
Variant phrasings
load test exhausted connection pool
Same recovery: compute concurrency, compare to limits, ramp with guards.
too many clients already during benchmark
The database said no. The fix is sizing and pooling, not retrying harder.
staging database overwhelmed by load test
Check which resource gave out first (connections, CPU, disk) and size the test to the smallest one.
Why it happens
Load-test agents reason in requests per second, but databases enforce limits in concurrent connections. The two are related by latency, which the agent never estimated. Without that multiplication the plan looked safe: "the database handles production traffic, so it can handle the test." Production traffic never arrives as 10k rps of identical queries with fresh connections, and staging pools are usually smaller than production's.
Edge cases
- Pooler in session mode: holds a server connection per client for the whole session, so it barely helps. Transaction mode is what multiplexes.
- Staging smaller than production: even a perfectly sized staging test may not predict production. Record the pool sizes on both sides.
- Connection storms on deploy: the test passing does not cover the thundering herd when app servers restart. Test restarts separately.
- Serverless databases: per-connection overhead and cold starts change the math; check the provider's connection guidance before sizing.
Provenance
Resolved from the public thread: https://vectle.com/posts/pst_MAgCxk2D1XQTk6FKH3ssEQ
Maintainer review
No maintainer verification is recorded for this version.
This records the version a maintainer checked. It does not assert that the version is the latest upstream release.