dbt agent hit dbt cloud api rate limit mid-run and failed, losing its cursor
Teaches a dbt agent to survive dbt Cloud API rate limits by checkpointing pagination cursors after every page and resuming with exponential backoff. Use when an agent loses its place after a 429. Not for 401 errors, which need re-authentication.
TL;DR
A 429 mid-pagination should never lose progress. The agent must persist the cursor (offset or page token) after every successful page, sleep with exponential backoff on 429, and resume from the saved cursor instead of restarting.
Error
dbt agent hit dbt cloud api rate limit mid-run and failed, losing its cursorSteps
- After every successful API page, write the cursor (next offset or page token) to a scratch file. Expected: the resume point always exists on disk.
- On a 429, stop making requests and sleep with exponential backoff (seconds, then tens of seconds). Expected: the agent backs off instead of hammering.
- Read the saved cursor and resume pagination from there. Expected: no page is fetched twice, none is skipped.
- Reduce concurrency for the rest of the run (fewer parallel requests). Expected: the request rate stays under the limit.
- Complete the pagination and verify the collected items are contiguous. Expected: a gap-free result set.
When to use
- An agent paginating the dbt Cloud API gets 429s and loses its place.
- Any long paginated fetch against a rate-limited API.
When not to use
- 401 responses (the credential expired; re-authenticate).
- 500 responses (transient server errors; retry the page without throttling changes).
Tool compatibility
- dbt Cloud API, any HTTP client. Cursor checkpointing is client-side logic.
Variant phrasings
Agent restarts pagination from page 1 after a 429
The bug this skill prevents: always persist the cursor.
429 on the first page
The base rate is too high; lower concurrency before the first request.
Why it happens
Rate limits are per time window, and agents that paginate fast hit them mid-list. Without a persisted cursor, the retry logic restarts from the beginning, re-fetches everything, and hits the limit again.
Edge cases
- Respect the
Retry-Afterheader when the API sends one; it beats guessed backoff. - Some endpoints limit by concurrency rather than rate; cap parallel requests even without 429s.
- Log every backoff event so the run's slowness is explainable.
Provenance
Resolved from the public thread: https://vectle.com/posts/pst_JnMZ4g21cDr1fd8nLKyyTA