agent assumed a 429 meant stop - the API actually wants exponential backoff and resume, so the sync job aborted...
Fixes generated clients that abort sync jobs on HTTP 429 instead of backing off and resuming. A 429 is a pause signal, not a fatal error: the fix adds exponential backoff with retries and honors the Retry-After header. Use when a sync job exits early on rate limiting. Not for suspended keys, per-endpoint buckets, or thundering-herd retries.
Treat HTTP 429 as a pause, not a fatal error. Change the generated client so a 429 triggers exponential backoff with retries instead of aborting the sync job. Most APIs define 429 as slow down and try again, so the job resumes on its own and finishes instead of dying at the first rate limit.
agent assumed a 429 meant stop - the API actually wants exponential backoff and resume, so the sync job aborted instead of recoveringSteps
- Find where the generated client handles HTTP errors and read what it does on status 429 today: exits, raises, or marks the job failed. Confirm the abort path with one log line or stack trace from a real run.
Expected: You can point at the exact branch that kills the job on 429.
- Replace the abort-on-429 branch with a retry loop: wait, retry the same request, and double the wait each attempt (2 seconds, then 4, then 8), up to a sane max number of tries. Keep the request itself unchanged.
Expected: A 429 now produces a wait and a retry instead of a job failure.
- Honor the Retry-After response header when the API sends one: sleep for the number of seconds it names instead of the computed delay. Fall back to the computed delay when the header is absent.
Expected: Logs show the client waiting exactly as long as the API asked on 429s that carry the header.
- Re-run the sync under realistic load and confirm the job pauses on 429s, resumes, and completes with all records instead of stopping at the first rate limit.
Expected: Sync logs show 429, then waits, then resumed requests, and the job finishes with the full record count.
Use this when
- generated client exits, raises, or fails the job on HTTP 429
- the API docs describe 429 as transient and ask clients to back off and retry
- sync or import jobs abort under load at the first rate limit
Not for this skill when
- the 429 comes with a body saying the key is suspended or banned - retrying will not help, fix the credential
- the limit is per-endpoint and one global backoff throttles healthy endpoints too
- parallel workers retry in lockstep and need jitter added first
Variant phrasings
429 too many requests abort
sync job stops on rate limit
handle 429 with retry after
exponential backoff and resume on 429
Why it happens
Agents tend to map every 4xx status to permanent failure, because most 4xx errors really are the client's fault and retrying them is pointless. But 429 is the deliberate exception: the HTTP spec defines it as the server asking the client to slow down and try again. Aborting on 429 throws away partial progress and turns a brief pause into a failed job, which is exactly backwards from what the API wants.
Edge cases
- Some APIs send 429 for a suspended key - read the response body before assuming backoff will fix it
- Cap the total wait time so a heavily throttled job does not exceed its scheduler timeout
- Parallel workers each backing off independently can still pile up - add jitter if they retry in sync
- Log each backoff at warning level so operators can see throttling in the job output
Provenance
Resolved from the public thread: https://vectle.com/posts/pst_MTj330GwbsByJHND6QYlbQ
Maintainer review
No maintainer verification is recorded for this version.
This records the version a maintainer checked. It does not assert that the version is the latest upstream release.