# Fireworks rate limits are tokens per minute, adaptive, and 503 can hit you anyway Three things about Fireworks serverless rate limits that trip people: 1. Enforcement is TPM, not TPS. Tokens per minute, broken into Total Prompt TPM (cached + uncached), Uncached Prompt TPM, and Generated TPM. A tokens-per-second budget will mislead you. 2. The limits are adaptive: they grow and shrink with your usage within ceilings set by model size (small models get higher ceilings; unknown parameter counts default to the large-model ceiling). Ramping traffic too fast gets you 429s even if your steady-state volume is fine. 3. Staying under your rate limit does not guarantee success. When a deployment is busy your traffic can be load shed with 503 Service Overloaded. Priority tier decreases the chance but does not remove it. Operational takeaways: back off exponentially on 429, treat 503 as a load-shedding signal not a bug, and remember limits are per account and per model with Fast and regular variants counted separately. And if your launch traffic exceeds the starting limit on day one, contact Fireworks for a custom ceiling instead of trying to warm the adaptive ramp with real users.

Context: Fireworks serverless rate limits: adaptive, TPM not TPS, 503 load shedding inside limits Fireworks serverless rate limits are adaptive: ceilings depend on the model's total parameter count (Small [400B gets 64.8M total prompt TPM; Large ]=1.6T gets 21.6M; unknown counts default to Large). Enforcement uses TPM, not TPS. Your effective limits grow and shrink with usage, and ramping traffic too quickly produces 429s. Staying under the limit does not guarantee success: busy deployments load-shed with 503 Service Overloaded, and Priority tier only reduces the chance. Limits are per account and per model; Fast and regular variants have separate limits.