how to debug RDS failover connectivity gaps
Diagnoses the seconds-to-minutes where apps cannot reach RDS during a failover. Use when Multi-AZ failovers cause longer outages than expected, when apps do not reconnect, or when DNS caching defeats the failover endpoint. Not for setting up Multi-AZ or for query performance.
TL;DR
RDS failover gaps are usually not RDS being slow; they are the app holding onto dead connections or cached DNS. The fix is on the client side: short DNS TTLs respected by the runtime, aggressive connection validation, and retry logic that actually reconnects. Measure the gap from the app's perspective, because RDS metrics will tell you failover was fast while your app was still down.
The query
how to debug RDS failover connectivity gapsUse this when
- Multi-AZ failovers cause longer app outages than the RDS event suggests
- Apps throw connection errors for minutes after a failover completes
- You are testing failover and the numbers do not add up
- Connection pools seem to hold dead connections
Not for when
- Enabling Multi-AZ for the first time
- Slow queries or database performance tuning
- Aurora-specific failover (similar ideas, different mechanics)
Steps
Step 1: Measure the gap from the application's side
Log connection errors with timestamps during a failover test. Compare the app's error window to the RDS event timeline. The difference between "RDS says failover took 60 seconds" and "the app errored for 4 minutes" is your real problem, and it lives client-side. Expected output: two timelines showing the RDS failover duration vs the app's outage duration.
Step 2: Check DNS caching in the app runtime
RDS failover works by flipping the DNS endpoint to the standby. If your runtime caches DNS (JVMs are famous for this, Node and Go have their own behaviors), the app keeps dialing the old primary. Verify the TTL your app actually honors, not just the DNS record's TTL. Expected output: confirmation of whether the app resolves the endpoint fresh on reconnect, or a found cache to disable/shorten.
Step 3: Validate connections before use
Configure the connection pool to test connections before handing them out (validation query, short validation timeout). After a failover, pools are full of dead connections; validation evicts them instead of handing failures to your code. Expected output: post-failover, the first requests after reconnect succeed instead of burning through dead pooled connections.
Step 4: Make retries reconnect, not just resend
Ensure retry logic creates a new connection (and re-resolves DNS) rather than retrying on the same broken socket. Set socket timeouts so a dead connection fails fast instead of hanging until TCP gives up. Expected output: retries during the failover window either succeed on the new primary or fail fast with a clear error, never hang.
Step 5: Run a real failover test and record the numbers
Trigger a failover (reboot with failover on a test instance) and measure the app-visible gap with the fixes in place. Write down the number; it is your actual RTO for the database, and it belongs in the runbook. Expected output: a measured, repeatable failover gap (often 60-120 seconds with proper client config) documented for the team.
Provenance
Resolved from the public thread: https://vectle.com/posts/psttYlnvRMLJxqGqGaNZaiww
Maintainer review
No maintainer verification is recorded for this version.
This records the version a maintainer checked. It does not assert that the version is the latest upstream release.