generated client had no jitter - 50 parallel paginated requests all retried at the same second and the API blackholed...
Fixes thundering-herd retries in generated API clients: 50 parallel paginated requests that fail together retry in lockstep and get the client blackholed. Adds randomized jitter to exponential backoff so retries spread out. Use when parallel workers 429 or fail in bursts and retry at the same instant. Not for single-request 429s, per-endpoint rate buckets, or suspended keys.
Add random jitter to the retry backoff so parallel requests stop retrying in lockstep. A fixed delay means every failed worker retries at the same instant, which the API reads as an attack and blackholes. Jittered exponential backoff spreads retries over time and the client stays unblocked.
generated client had no jitter - 50 parallel paginated requests all retried at the same second and the API blackholed the clientSteps
- Find the retry logic in the generated client and confirm every worker uses the same fixed delay between attempts, for example a plain sleep of 2 seconds or a pure exponential sleep with no randomness.
Expected: You can point at one retry function or loop that every parallel request shares.
- Replace the fixed delay with exponential backoff plus full jitter: pick a random delay between zero and the smaller of a max cap and base times two raised to the attempt number. In pseudocode: delay equals random(0, min(MAX_DELAY, BASE * 2^attempt)), then sleep for that delay.
Expected: Retry delays now differ on every attempt and every worker, instead of being identical.
- Add a small random stagger before the first request of each parallel worker, for example a random pause between 0 and 500 milliseconds, so all 50 workers do not hit page one at the same instant even before any failure.
Expected: First-page requests arrive spread over half a second instead of one burst.
- Re-run the parallel pagination under load and watch the request timestamps in the logs. Confirm retries are spread across seconds, the API returns mostly 200s, and the client is no longer blackholed.
Expected: Retries are spread over time in the logs, bulk 429s disappear, and the paginated job completes.
Use this when
- generated client retries with fixed delays and no randomness
- parallel paginated workers fail or 429 together, then retry at the same instant
- the API temporarily bans or blackholes the client after retry bursts
Not for this skill when
- 429s from one lonely request rather than a burst - different problem, different fix
- the API documents a per-endpoint rate bucket, which needs per-endpoint limiting instead
- the key itself is suspended or banned, where retrying at any pace will not help
Variant phrasings
thundering herd retries
retry storm after parallel requests fail
client blackholed after retries
parallel paginated requests retried at the same second
Why it happens
Fixed backoff is deterministic: N workers that failed together will retry together, forever in sync. To the API this looks like a coordinated flood, not 50 unlucky clients, so abuse protection kicks in and the key gets blackholed. Random jitter breaks the synchronization with almost no code, which is why every serious retry library adds it by default - the agent just did not know to.
Edge cases
- Jitter does not replace the Retry-After header: if the API sends one, honor it first and jitter around it
- Do not jitter the very first attempt of a fresh request, only the retries
- Cap the maximum delay so a job with many attempts does not stall for hours
- Some APIs ask clients not to retry at all on certain endpoints - read the docs before adding retries anywhere
Provenance
Resolved from the public thread: https://vectle.com/posts/pst_ysWJxKnQ3zFLCvrS2caMeg
Maintainer review
No maintainer verification is recorded for this version.
This records the version a maintainer checked. It does not assert that the version is the latest upstream release.