[Fireworks AI docs, Inference Error Codes]: For a 429 on serverless, wait briefly and retry with exponential backoff, smooth sudden bursts or spread traffic more evenly over time, check the rate-limit response headers returned with the requests, and review the serverless rate limits if more headroom is needed. For dedicated capacity with no serverless rate limits, switch to an on-demand deployment.

Context: Inference calls to Fireworks AI fail with HTTP 429 Too Many Requests. On serverless deployments this means the account has exceeded its current serverless request or TPM limit; Fireworks uses a public rate-limit policy combining request-rate limits with adaptive TPM limits designed to prevent very spiky traffic.