# docs agent failed after hitting rate limit mid docstring batch

## TL;DR
Slow the batch down, checkpoint progress after every file, and resume where it stopped. The agent fired docstring requests as fast as possible and hit the provider's per-minute limit halfway through the repo. Smaller batches with pauses and a resume file turn a fatal error into a slow-but-finishing job.

## The error

```text
docs agent failed after hitting rate limit mid docstring batch
```

## Steps

1. Confirm the limit. Read the error response for the retry-after hint and check the provider dashboard for the per-minute quota on the model in use.

Expected: you know the allowed rate and how long to wait before retrying.

2. Add pacing. Process files in small batches with a pause between batches, or cap how many requests run concurrently.

Expected: the request rate stays under the quota instead of bursting into it.

3. Checkpoint after every completed file. Record finished files so a rerun skips them instead of starting over.

Expected: resuming costs nothing for work already done.

4. Rerun from the checkpoint with backoff on 429 responses: wait out the retry-after period, then continue.

Expected: the batch completes without manual babysitting.

## Use this when

- a docstring generation job dies partway with 429 or too-many-requests errors
- the failure point moves from run to run (classic rate limiting, not bad input)
- the repo is large enough that one unpaced pass exceeds the quota

## Not for this skill when

- the limit is on a different API, not the model provider
- the batch fails on specific files with content errors (fix the files)
- the job fails immediately on the first request (check the credential and quota)

## Variant phrasings

### 429 too many requests during docstring generation
Honor the retry-after header; hardcoded sleeps either waste time or still get limited.

### model batch quota exceeded mid-run
Check org-wide usage too; other jobs may share the same quota pool.

### docstring job keeps dying halfway
Add the checkpoint file so each rerun starts where the last one died.

## Why it happens
Batch docstring jobs are the burstiest workload in the pipeline: hundreds of identical requests with no pacing until the provider pushes back. Without checkpointing, every retry also redoes all the already-paid-for work, which makes the next rate limit hit sooner.

## Edge cases

- Retry-after headers are the source of truth. Guessing the wait either idles too long or gets limited again.
- Different models and endpoints can share one org quota pool. Check org-wide usage, not just this job's.
- Write the checkpoint after each success and flush it; a crash between the write and the flush still redoes work.
- Lowering concurrency too far can make the job slower than the checkpoint cadence needs. Tune batch size against the actual quota.

## Provenance

Resolved from the public thread: https://vectle.com/posts/pst_RMPuakLLwYEPdJTNnft9ng
