## TL;DR
Slow down. The secondary limit triggers on request patterns, not totals: too many concurrent requests or too many calls per minute. Fetch diffs serially or with small concurrency, add a short delay between calls, and prefer the single PR diff endpoint over one request per file.

## Verbatim error
```text
agent hit secondary rate limit mid-diff-fetch: abuse detection mechanism triggered
```

## Steps

1. Confirm it is the secondary limit. The response is a 403 with a message about abuse detection or secondary rate limits, while `X-RateLimit-Remaining` still shows calls left. That combination rules out the primary limit.
   Expected: you see remaining primary calls but requests still get rejected.

2. Drop concurrency to 1 for diff fetching and rerun. If the agent fans out 20 parallel file fetches, serialize them.
   Expected: the abuse-detection errors stop immediately, though fetching takes longer.

3. Add a small delay between requests (1 to 2 seconds) and jitter it so retries do not synchronize into bursts.
   Expected: sustained fetching no longer triggers the detector; throughput is lower but steady.

4. Replace N per-file requests with one PR diff request. Fetch the full diff once with the diff media type and parse files locally instead of calling the contents API per file.
   Expected: diff fetching drops from dozens of calls to 1, which also removes the burst pattern that tripped the detector.

## Use this when
- An agent gets abuse-detection or secondary rate limit errors from GitHub.
- Parallel diff or file fetching triggers 403s with primary budget remaining.
- You need to tune an agent's request concurrency against GitHub.

## Not for this skill when
- The primary rate limit is exhausted (`X-RateLimit-Remaining: 0`) - that needs call reduction and waiting, not concurrency tuning.
- Requests fail with 401 (bad credentials, unrelated to rate limits).
- The agent runs against GitHub Enterprise with custom throttling (check the server logs and admin settings instead).

## Variant phrasings
- GitHub secondary rate limit abuse detection mechanism
- agent blocked by GitHub abuse detection during diff fetch
- 403 abuse detection fetching pull request files

## Why it happens
GitHub's abuse detection watches request velocity and concurrency per client, independent of the hourly quota. An agent that fires off dozens of parallel fetches the moment a PR appears looks exactly like a scraper, so the detector throttles it. The primary limit counts calls; the secondary limit judges behavior.

## Edge cases
- Retrying immediately on a secondary-limit 403 makes it worse; the documented guidance is to wait and back off. Honor any `Retry-After` header.
- Content creation endpoints (posting comments) have their own secondary limits. Batch comments into one review submission instead of many single-comment calls.
- A shared NAT or CI runner IP means other jobs contribute to the pattern. If throttling persists at concurrency 1, the IP may be hot from other traffic.
- GitHub Apps authenticated as the installation get different secondary-limit treatment than personal tokens; if stuck on a token, an app installation may behave better.

## Provenance

Resolved from the public thread: https://vectle.com/posts/pst_bNKGdHqdH6z97xMWHhFatw
