## TL;DR
A flapping VPN tunnel is almost always DPD timeouts, rekey mismatches, MTU blackholes, or one side restarting. Work the checklist top to bottom: confirm both ends see the same state, check lifetimes and DPD settings match, then look at packet size and provider-side health. Fix the config mismatch, not the tunnel.

## Error / query
```text
VPN tunnel flapping: troubleshooting checklist
```

## Use this skill when
- a site-to-site VPN goes up and down every few minutes or hours
- one AWS VPN tunnel is up while the other flaps
- traffic drops intermittently across an IPsec tunnel
- you just changed a firewall or VPN config and the tunnel got unstable

## Not for this skill when
- the tunnel has never come up at all (check proposals and PSKs first)
- the issue is throughput, not stability
- you are debugging client VPN (OpenVPN/WireGuard road warrior) rather than site-to-site

## Steps

### Step 1: Confirm the flap pattern and which side initiates the drop
```bash
aws ec2 describe-vpn-connections --vpn-connection-ids [VPN_ID] --query "VpnConnections[0].VgwTelemetry"
journalctl -u strongswan --since "2 hours ago" | grep -i -E "rekey|dpd|delet|establish" | tail -30
```
Expected: telemetry shows UP/DOWN transitions with timestamps; the daemon log shows whether your side or the peer tore the SA down (look for DPD timeout vs delete notify).

### Step 2: Check SA lifetimes and rekey settings on both ends
```bash
swanctl --list-sas --raw | grep -E "rekey|life" | head -20
ipsec statusall | grep -E "life|rekey" | head -20
```
Expected: IKE and ESP lifetimes printed for each SA. Both ends should agree within a small margin; a big mismatch means one side rekeys while the other still uses the old SA, which reads as a flap.

### Step 3: Verify DPD is enabled and aggressive enough
```bash
grep -r -i "dpd" /etc/strongswan.d/ /etc/ipsec.conf | head -10
```
Expected: dpdaction and dpddelay settings visible. If DPD is off or the delay is longer than the peer's idle timeout, dead tunnels linger and then drop all at once. Enable DPD with a delay shorter than any NAT or firewall idle timeout on the path.

### Step 4: Rule out MTU and MSS blackholes
```bash
ping -M do -s 1400 -c 5 [PEER_INSIDE_IP]
ping -M do -s 1200 -c 5 [PEER_INSIDE_IP]
```
Expected: the large packet fails and the smaller succeeds if there is an MTU blackhole. IPsec overhead eats 60 to 80 bytes; clamp MSS on the tunnel interface so TCP sessions do not hang mid-transfer.

### Step 5: Check provider-side health and failover behavior
```bash
aws cloudwatch get-metric-statistics --namespace AWS/VPN --metric-name TunnelState --dimensions Name=TunnelIpAddress,Value=[TUNNEL_IP] --start-time [START] --end-time [END] --period 300 --statistics Maximum
```
Expected: a per-tunnel time series. If your side shows UP while the provider shows DOWN, the problem is usually on your side's config; if both flap together, look at the path (BGP, ISP, or the peer device rebooting).

## Variant phrasings

### "ipsec tunnel drops every hour"
Classic lifetime or rekey mismatch. Compare the IKE/ESP lifetimes on both ends first (step 2); an hourly flap lines up with a 3600s lifetime somewhere.

### "aws vpn one tunnel up one tunnel down"
AWS gives you two tunnels for redundancy and they flap independently. Check the per-tunnel telemetry and make sure your device answers DPD on both tunnel IPs, not just the first.

### "vpn stable but traffic stalls randomly"
That is the MTU symptom, not a flap. Run step 4 and clamp MSS; the tunnel staying up while big packets die is the giveaway.

## Why it happens
IPsec tunnels are two independent security associations (IKE and ESP) that both sides must keep in sync. Anything that breaks the sync looks like flapping: mismatched lifetimes causing unsynchronized rekeys, DPD gaps letting dead SAs linger, NAT or firewall idle timeouts killing the UDP 500/4500 flow, or MTU issues that only affect real traffic and not keepalives.

## Edge cases and pitfalls
- Both tunnels flapping at the same time usually means the peer device or the path, not your config; check BGP session state too.
- Some peers ignore DPD entirely; if the far end never answers DPD, tune your side to restart on timeout rather than waiting.
- Rekey collisions happen when both ends initiate at once; a small random fuzz on lifetimes avoids the thundering rekey.
- After any firewall change, re-verify UDP 500 and 4500 plus ESP (protocol 50) are open in both directions.
- Log retention matters: enable IKE debug logging before the next flap, because post-mortem without logs is guessing.

## Provenance

Resolved from the public thread: https://vectle.com/posts/pst_9JapZG2wRzPl7CYzjpaeCw
