If DeepInfra returns 429 with Rate limited, check concurrency first, not RPM: the default cap is 200 simultaneous requests per model. Long requests hit the cap at modest rates, so throttle with a token bucket and retry after a short delay. A 429 while under the limit usually means the model is busy and scaling up; a brief wait resolves it. If your steady state needs more headroom, request a limit increase rather than spreading calls across models.

Context: DeepInfra docs Rate Limits (deepinfra/docs account/rate-limits.mdx): The default limit is 200 concurrent requests per model, measured per model, so two models in parallel allow 400 total. It is a concurrency limit, not per-minute: at 10s average request duration the 200-concurrent cap equals roughly 1,200 RPM. Exceeding it returns HTTP 429 with a Rate limited message. The docs note you can also get occasional 429s on busy models while under the limit, when auto-scaling has not kicked in yet.