anthropic overloaded_error: Overloaded
Handles Anthropic overloaded_error: Overloaded responses. Use when Claude API calls fail with overloaded, when an agent's pipeline needs resilience, or during known capacity incidents. Not for auth errors, for invalid model names, or for context-length errors.
TL;DR
overloaded_error means Anthropic's serving capacity is saturated right now; it is transient and the correct response is retry with exponential backoff, not a code fix. Implement retries with jitter, shed non-critical load, and check the status page to distinguish an incident from normal contention.
anthropic overloaded_error: OverloadedUse this when
- Claude API calls fail with overloaded_error
- An agent pipeline needs to survive capacity blips
- Errors cluster in time then clear on their own
Not for this skill when
- The key is invalid (thats auth)
- The model name is wrong (thats not-found)
- Prompts exceed context (thats a 400, not overload)
Steps
- Confirm it is transient, not your bug:
Check https://status.anthropic.com for an active incidentExpected output: either an incident (wait it out) or all-clear (your retries will succeed). Overload is never fixed by changing your prompt.
- Add exponential backoff with jitter around the call:
import time, random
for attempt in range(5):
try:
return call_claude()
except OverloadedError:
time.sleep((2 ** attempt) + random.random())
raise RuntimeError("still overloaded after retries")Expected output: most blips clear within a few retries. Jitter keeps your retries from stampeding with everyone else's.
- Reduce concurrent load during the incident:
# lower worker count / add a semaphore around API callsExpected output: fewer in-flight requests means fewer overload errors. An agent fleet hammering the API during an incident makes its own problem worse.
- For pipelines that cant wait, queue and defer:
# push failed items to a retry queue with a delay instead of failing the runExpected output: the pipeline completes late rather than failing. Batch and offline workloads should treat overload as "try later," not "error."
Variant phrasings
overloaded only at certain hours
Peak contention. Shift batch work off-peak or accept retries as normal during those windows.
529 vs overloaded_error
Same family: 529 is the HTTP status some clients surface for the same condition. Same handling.
Why it happens
LLM serving has finite GPU capacity; when demand spikes (new model launch, big incident reroutes, your own fleet scaling up), the API sheds load with overloaded_error instead of queueing forever. It is a backpressure signal, and the API contract is "retry later." Treating it as a fatal error is the bug in your code, not theirs.
Edge cases
- Dont retry instantly in a tight loop; that is a self-inflicted DDoS and prolongs the incident for you.
- Long-running agents should distinguish overload (retry) from auth/billing errors (stop immediately).
- If overload persists for hours, check whether your account or tier changed; sustained overload can also reflect reduced quota.
Provenance
Resolved from the public thread: https://vectle.com/posts/pstdEh0ioimG0G5Ply4Ts0bQ
Maintainer review
No maintainer verification is recorded for this version.
This records the version a maintainer checked. It does not assert that the version is the latest upstream release.