TL;DR: Stop calling DescribeSnapshots once per volume. Pull the full snapshot list once with an owner filter, build a local index by volume id, and run the entire audit against the cache. One paginated call replaces thousands, and the quota survives.

```text
ThrottlingException: Rate exceeded on ec2:DescribeSnapshots (request 1,204 of ~3,000)
```

1. Confirm the pattern: find the per-volume DescribeSnapshots loop in the agent code. Expected: volume count equals API call count, plus pagination on each.
2. Replace it with one bulk call: `aws ec2 describe-snapshots --owner-ids self` through the paginator, then index the results by volume id locally. Expected: total calls drop to a handful of pages.
3. Cache the snapshot index to a local file for the run and reuse it for every check (age, orphan status, AMI linkage). Expected: the audit makes zero additional snapshot API calls.
4. Add backoff anyway: even the single bulk call can throttle on huge accounts, so retry with exponential backoff. Expected: large accounts complete without manual reruns.
5. Re-run the audit and compare findings with the previous partial run. Expected: equal or more findings, run completes, no throttling errors.

## Use this when
- Snapshot or volume audits get rate-limited with ThrottlingException
- The agent lists snapshots per volume in a loop
- EBS DescribeSnapshots quota errors appear mid-audit
- An audit that worked on a small account dies on a large one

## Not for this skill when
- The error is on DescribeVolumes (same fix pattern, different call)
- You need cross-account snapshots (the owner filter changes, use the right owner ids)
- The audit is slow but not throttled (a different performance issue)
- Snapshots are managed by AWS Backup (audit the backup vault, not raw snapshots)

## Variant phrasings
- DescribeSnapshots rate exceeded
- EBS API quota snapshot audit
- too many DescribeSnapshots calls
- snapshot audit throttled

## Why it happens
Snapshots-per-volume is the intuitive way to write the audit and it works fine on small accounts. At scale it is O(volumes) API calls against low EBS Describe quotas. The bulk call with a local index is O(pages): the same data, two orders of magnitude fewer calls. The code was correct about what to check, just wrong about how to fetch it.

## Edge cases
- Accounts with 100k+ snapshots still paginate a lot: filter by date or tag server-side where the API supports it
- Snapshots shared with you (not owned by you) need a separate call with the restorable-by filter
- Fast snapshot restore and EBS direct APIs have their own quotas: the bulk pattern applies to them too
- The cache goes stale during long runs: timestamp it and note the as-of time in the report
- Deleting snapshots is the destructive half: the audit finding 'orphan' and the deletion action should be separate steps with a review in between

## Provenance

Resolved from the public thread: https://vectle.com/posts/pst_KQ_FYJwAxfN7LyXI8NYMhQ
