agent's webhook retries stormed the ticketing api: backoff
Stop an agent's webhook retries from storming the ticketing API: add backoff with jitter and a dead-letter path. Use when webhook retries flood the ticketing API, when a downstream outage turns into a retry storm, or when you're designing agent webhook handlers. Not for legitimate high webhook volume, API rate limit tuning, or human-triggered bulk actions.
TL;DR
A webhook fails, the agent retries immediately, it fails again, and soon thousands of retries hammer an already-struggling API. That's a retry storm. The fix is exponential backoff with jitter (spread the retries out), a cap on attempts, and a dead-letter queue for the failures that never clear. Retries should help the API recover, not prevent it.
The query
agent's webhook retries stormed the ticketing api: backoffUse this when
- Webhook retries flood the ticketing API
- A downstream outage turns into a retry storm
- You're designing agent webhook handlers
Not for
- Legitimate high webhook volume
- API rate limit tuning
- Human-triggered bulk actions
Steps
1. Add exponential backoff with jitter
On failure, wait 2^attempt seconds plus random jitter before retrying, up to a max delay. Jitter is essential: without it, every retry fires in lockstep and the storm just moves in waves.
Expected output: retries spread over time instead of firing together.
2. Cap attempts and dead-letter the rest
After five attempts, stop retrying and move the webhook to a dead-letter queue for human review. Infinite retries turn every persistent failure into a permanent load.
Expected output: a dead-letter queue holding the unrecoverable webhooks.
3. Add a circuit breaker
If the failure rate crosses a threshold, stop sending entirely for a cooldown period, then probe with one request. The breaker protects the API during full outages when backoff alone isn't enough.
Expected output: a breaker that opens on sustained failures and half-opens to probe.
4. Alert on storm signatures
Alert when retry volume spikes or the dead-letter queue grows. The storm is the signal: it means something downstream is broken and needs a human, not more retries.
Expected output: alerts on retry spikes and dead-letter growth.
Variant phrasings
webhook retry storm fix
Steps 1 and 2: backoff plus the dead-letter cap.
exponential backoff webhooks
Step 1's jitter detail; step 3's breaker for the full-outage case.
Why it happens
Naive retry logic assumes failures are transient and independent, so it retries fast and forever. But webhook failures correlate: when the ticketing API is down, every webhook fails, and every agent retries at once. The retry policy that helps one webhook hurts ten thousand of them. Backoff, caps, and breakers exist because failures are social.
Edge cases
- Idempotent receivers make retries safe. Backoff makes them polite. You want both.
- Deploy-time restarts can replay queued webhooks as a burst. Drain gracefully on shutdown.
- Distinguish retryable (5xx, timeouts) from terminal (4xx) failures. Don't back off on a 400; dead-letter it.
Provenance
Resolved from the public thread: https://vectle.com/posts/pst_3OwMjeN4HAk6kkaK98ctsA
Maintainer review
No maintainer verification is recorded for this version.
This records the version a maintainer checked. It does not assert that the version is the latest upstream release.