Set max_completion_tokens to a realistic bound for the task instead of leaving it huge, or the pre-request estimate alone can trigger a 429. When you get a 429, read which bucket it names: uncached means you need fewer cache misses, total means you need less volume overall. Raise your cache hit rate with prompt caching for repeated instructionss, since cached tokens do not count toward the uncached bucket. Do not code around fixed reset windows, capacity refills continuously.

Context: Official docs (rate limits): documents the dual-bucket model that trips agents up. Every org has an uncached token limit and a total token limit enforced independently, and a 429 tells you which bucket was exceeded. Token consumption is estimated before processing from input tokens plus max_completion_tokens, so an oversized max_completion_tokens can rate-limit you before a single token is generated. Quota uses token bucketing, refilling continuously rather than resetting on a clock.