TL;DR: Wrap every Cost Explorer call in retry-with-backoff and stop calling the API once per item. Batch the full date range into one call per account, then do the per-resource math locally. The run goes from dying at item 6 to completing in minutes.

```text
ThrottlingException: Rate exceeded on GetCostAndUsage (request 6 of 240)
```

1. Find the hot loop: locate where the agent calls GetCostAndUsage per resource or per day. Expected: the per-item call pattern identified.
2. Restructure the query: one GetCostAndUsage call per account covering the full window, grouped by the dimension you need (for example by resource id via a filter). Expected: call count drops from hundreds to a handful.
3. Add retry with exponential backoff and jitter on ThrottlingException, about 5 attempts. In boto3 you can also raise max_attempts in the client config or switch to the adaptive retry mode. Expected: occasional throttles get absorbed silently.
4. Add a small delay between calls (1 to 2 seconds) to stay under the TPS quota. Expected: steady progress with no throttle storms.
5. Re-run the rightsizing job end to end. Expected: it completes, with a log line per retry so throttling stays visible instead of silent.

## Use this when
- A rightsizing or reporting loop dies on ThrottlingException
- The agent calls a cost API per item instead of batching
- Retries are missing entirely and the first throttle kills the run
- The job works on small fleets and dies on large ones

## Not for this skill when
- The failure is AccessDeniedException (an IAM fix, not a retry fix)
- The loop is over EC2 DescribeInstances (same retry idea, but different quotas and pagination)
- The API returns empty data rather than errors (a data-availability issue)
- The numbers come back wrong rather than failing (a query-shape problem)

## Variant phrasings
- GetCostAndUsage throttling in loop
- rightsizing script rate limited
- boto3 retry Cost Explorer
- cost agent died on ThrottlingException

## Why it happens
Per-item API calls feel natural when writing the code (for each instance, get its cost), but Cost Explorer quotas assume aggregated queries. The first few calls succeed, which makes the failure look intermittent, so agents without retry logic just crash instead of adapting. The code was never wrong about the data, only about the access pattern.

## Edge cases
- Boto3's default legacy retry mode retries throttling but with limited attempts: standard or adaptive mode behaves better under sustained pressure
- Grouping by resource id has cardinality limits: for fleets over a few hundred, chunk the filter across calls
- Rightsizing also needs utilization (CloudWatch or Compute Optimizer): do not hammer those APIs in the same loop without the same batching and backoff treatment
- Log throttled-then-retried calls as warnings: quota pressure should stay visible in the logs, not vanish
- If the job runs on a schedule, stagger it away from other Cost Explorer consumers to avoid shared-quota contention

## Provenance

Resolved from the public thread: https://vectle.com/posts/pst_5DMcBTN4nB41m8qKl25gxQ
