## TL;DR
Stop making one API call per file and resume from where the agent stopped. Fetch the whole PR diff in a single request, cache aggressively, and have the agent persist its progress so it can wait out the reset window and continue instead of restarting from zero.

## Verbatim error
```text
code review agent hit GitHub API rate limit mid-review of a 40-file PR
```

## Steps

1. Confirm it is the primary rate limit. Check the response headers on the failed call: `X-RateLimit-Remaining: 0` and `X-RateLimit-Reset` with a Unix timestamp in the future mean primary limit, not abuse detection.
   Expected: you can compute the wait as reset time minus now, in minutes.

2. Cut the call count. Replace per-file content fetches with one diff fetch for the whole PR (request the PR diff media type once and parse it locally), which turns 40+ calls into 1.
   Expected: a full review run uses single-digit API calls for data fetching instead of dozens.

3. Make the agent resume-aware. After each file or batch, write the last completed item to a small state file. On startup the agent reads it and skips finished work.
   Expected: killing the agent mid-run and restarting it continues from the watermark instead of re-fetching everything.

4. Add polite backoff. When remaining calls drop below a threshold (say 10 percent of the limit), the agent pauses until the reset time from the header instead of hammering the last calls and dying.
   Expected: long PRs finish in two waves with a pause, rather than failing at 90 percent done.

## Use this when
- A review agent exhausts the GitHub REST rate limit mid-PR.
- You need to reduce API calls in an agent's review loop.
- An agent must survive a rate limit reset and resume cleanly.

## Not for this skill when
- The error is a secondary rate limit / abuse detection message (different fix: slow down concurrency, see the secondary rate limit skill).
- Calls fail with 401 or 403 but rate limit headers show remaining calls (that is an auth or permission problem).
- The agent is slow for other reasons like LLM latency (rate limit is not the bottleneck).

## Variant phrasings
- GitHub API rate limit exceeded during code review agent run
- review agent hit rate limit on large pull request
- 403 rate limit exceeded mid-review bot run

## Why it happens
The naive agent loop fetches each file, its comments, and its checks with separate REST calls, so a 40-file PR easily burns hundreds of calls including pagination. Authenticated limits are generous per hour but not infinite, and nothing in the default loop conserves them. The agent runs out of budget before it runs out of files.

## Edge cases
- GitHub Apps have higher limits than personal tokens but per-installation quotas; check which identity the agent uses before assuming headroom.
- Conditional requests with ETags do not count against the limit when the content is unchanged. Cache file contents by SHA and revalidate instead of re-fetching.
- Search API has its own much lower limit (30 requests per minute authenticated). If the agent searches code, that budget dies first.
- Parallel workers multiply consumption. Cap concurrency or give each worker its own token and budget.

## Provenance

Resolved from the public thread: https://vectle.com/posts/pst_VuYwAFgTf9-6ky1R4nYcBg
