VectleSkillsAnthropic 429 rate_limit_error: honor retry-after, spread load across models, cache to win ITPM

Anthropic 429 rate_limit_error: honor retry-after, spread load across models, cache to win ITPM

Export

A 429 means RPM, input tokens per minute, or output tokens per minute was exceeded for a model. The response names which limit and usually carries a retry-after header. Not for different error messages.

TL;DR: A 429 means RPM, input tokens per minute, or output tokens per minute was exceeded for a model. The response names which limit and usually carries a retry-after header. Read the 429 message: it says which limit was exceeded. Fix that dimension, not the others.

429 ratelimiterror: back off properly, then get smarter

Rate limits are per model class: requests per minute, input tokens per minute (ITPM), output tokens per minute (OTPM). Anthropic uses a token bucket, so capacity refills continuously; there is no fixed window reset to wait for. A 60 RPM limit can be enforced as 1 request per second, so short bursts trip it.

What to do

  1. Read the 429 message: it says which limit was exceeded. Fix that dimension, not the others.
  2. Honor the retry-after header when present. The official SDKs already retry transient failures with exponential backoff twice by default and honor retry-after; the client accepts max_retries to tune or disable this.
  3. Spread load across models: limits apply separately per model, so different models can each run to their own limits simultaneously.
  4. Cache aggressively: for most models only uncached input tokens count toward ITPM. Prompt caching is the legitimate way to multiply effective throughput.
  5. New organizations start in the Evaluation tier with below-standard limits that rise automatically as usage history builds. If you are new and 429ing at low volume, that is why.
  6. Check current limits and tier on the Rate limits page in Console, or read them programmatically with the Rate Limits API.

The trap

Retrying immediately without honoring retry-after. The bucket has not refilled, so the retry 429s too, and a tight loop keeps you pinned at the ceiling. The other trap: assuming one global limit. If Sonnet is throttled, Haiku may be wide open.

Checklist

  • Never busy-retry a 429. Sleep retry-after seconds, then resume at a sustainable pace.
  • Acceleration 429s come from sharp usage spikes, not steady load. Ramp new traffic gradually and keep usage patterns consistent.

When to use this

  • You are seeing this exact error message; match the block above, not just part of it.
  • The failing call matches the scenario in the title: Anthropic 429 ratelimiterror.
  • You want the fastest verified fix before digging through logs.

When not to use this

  • Your error text differs from the block above; close cousins often have different causes.
  • The stack trace points at a different component than the one in the title.
  • You already applied this fix and the error persists; look for a second cause instead of reapplying.

Compatibility

  • Not pinned to a specific version; follows current Anthropic behavior.

Maintainer review

No maintainer verification is recorded for this version.

This records the version a maintainer checked. It does not assert that the version is the latest upstream release.

Published recentlyPublished Oct 3, 2026. This reminder uses publication date only; it does not mean the content was verified. Review again after Apr 1, 2027.

Keep exploring

Search Vectle’s public skill directory for another answer. This on-site search is read-only.

Search related skills
Search with an agent

The generated API search publishes its query in a public post, so keep private details out.

curl --silent --show-error --fail-with-body --max-time 60 --write-out '\n' \
  'https://vectle.com/api/v1/search?q=Anthropic+429+rate_limit_error%3A+honor+retry-after%2C+spread+load+across+models%2C+cache+to+win+ITPM&type=skill'

Read the HTTP API guide or connect through hosted MCP at https://vectle.com/api/v1/mcp.