VPN tunnel flapping: troubleshooting checklist
Checklist for flapping site-to-site VPN tunnels: confirm the flap pattern and which side drops, compare SA lifetimes and rekey settings, verify DPD, rule out MTU blackholes, and check provider-side telemetry. Use when a tunnel cycles up and down or one AWS tunnel flaps while the other holds. Not for tunnels that never came up, throughput issues, or client VPN.
TL;DR
A flapping VPN tunnel is almost always DPD timeouts, rekey mismatches, MTU blackholes, or one side restarting. Work the checklist top to bottom: confirm both ends see the same state, check lifetimes and DPD settings match, then look at packet size and provider-side health. Fix the config mismatch, not the tunnel.
Error / query
VPN tunnel flapping: troubleshooting checklistUse this skill when
- a site-to-site VPN goes up and down every few minutes or hours
- one AWS VPN tunnel is up while the other flaps
- traffic drops intermittently across an IPsec tunnel
- you just changed a firewall or VPN config and the tunnel got unstable
Not for this skill when
- the tunnel has never come up at all (check proposals and PSKs first)
- the issue is throughput, not stability
- you are debugging client VPN (OpenVPN/WireGuard road warrior) rather than site-to-site
Steps
Step 1: Confirm the flap pattern and which side initiates the drop
aws ec2 describe-vpn-connections --vpn-connection-ids [VPN_ID] --query "VpnConnections[0].VgwTelemetry"
journalctl -u strongswan --since "2 hours ago" | grep -i -E "rekey|dpd|delet|establish" | tail -30Expected: telemetry shows UP/DOWN transitions with timestamps; the daemon log shows whether your side or the peer tore the SA down (look for DPD timeout vs delete notify).
Step 2: Check SA lifetimes and rekey settings on both ends
swanctl --list-sas --raw | grep -E "rekey|life" | head -20
ipsec statusall | grep -E "life|rekey" | head -20Expected: IKE and ESP lifetimes printed for each SA. Both ends should agree within a small margin; a big mismatch means one side rekeys while the other still uses the old SA, which reads as a flap.
Step 3: Verify DPD is enabled and aggressive enough
grep -r -i "dpd" /etc/strongswan.d/ /etc/ipsec.conf | head -10Expected: dpdaction and dpddelay settings visible. If DPD is off or the delay is longer than the peer's idle timeout, dead tunnels linger and then drop all at once. Enable DPD with a delay shorter than any NAT or firewall idle timeout on the path.
Step 4: Rule out MTU and MSS blackholes
ping -M do -s 1400 -c 5 [PEER_INSIDE_IP]
ping -M do -s 1200 -c 5 [PEER_INSIDE_IP]Expected: the large packet fails and the smaller succeeds if there is an MTU blackhole. IPsec overhead eats 60 to 80 bytes; clamp MSS on the tunnel interface so TCP sessions do not hang mid-transfer.
Step 5: Check provider-side health and failover behavior
aws cloudwatch get-metric-statistics --namespace AWS/VPN --metric-name TunnelState --dimensions Name=TunnelIpAddress,Value=[TUNNEL_IP] --start-time [START] --end-time [END] --period 300 --statistics MaximumExpected: a per-tunnel time series. If your side shows UP while the provider shows DOWN, the problem is usually on your side's config; if both flap together, look at the path (BGP, ISP, or the peer device rebooting).
Variant phrasings
"ipsec tunnel drops every hour"
Classic lifetime or rekey mismatch. Compare the IKE/ESP lifetimes on both ends first (step 2); an hourly flap lines up with a 3600s lifetime somewhere.
"aws vpn one tunnel up one tunnel down"
AWS gives you two tunnels for redundancy and they flap independently. Check the per-tunnel telemetry and make sure your device answers DPD on both tunnel IPs, not just the first.
"vpn stable but traffic stalls randomly"
That is the MTU symptom, not a flap. Run step 4 and clamp MSS; the tunnel staying up while big packets die is the giveaway.
Why it happens
IPsec tunnels are two independent security associations (IKE and ESP) that both sides must keep in sync. Anything that breaks the sync looks like flapping: mismatched lifetimes causing unsynchronized rekeys, DPD gaps letting dead SAs linger, NAT or firewall idle timeouts killing the UDP 500/4500 flow, or MTU issues that only affect real traffic and not keepalives.
Edge cases and pitfalls
- Both tunnels flapping at the same time usually means the peer device or the path, not your config; check BGP session state too.
- Some peers ignore DPD entirely; if the far end never answers DPD, tune your side to restart on timeout rather than waiting.
- Rekey collisions happen when both ends initiate at once; a small random fuzz on lifetimes avoids the thundering rekey.
- After any firewall change, re-verify UDP 500 and 4500 plus ESP (protocol 50) are open in both directions.
- Log retention matters: enable IKE debug logging before the next flap, because post-mortem without logs is guessing.
Provenance
Resolved from the public thread: https://vectle.com/posts/pst_9JapZG2wRzPl7CYzjpaeCw
Maintainer review
No maintainer verification is recorded for this version.
This records the version a maintainer checked. It does not assert that the version is the latest upstream release.