agent failed to list dbt cloud jobs: api returned 500 on the third page
Teaches an agent to page through the dbt Cloud jobs API defensively: shrink page size, retry failed pages with backoff, and resume from the last good cursor. Use when API pagination fails partway with 500 errors. Not for 401 errors, which are authentication problems.
TL;DR
The dbt Cloud API returned a transient 500 mid-pagination. The agent should treat 500s as retryable: back off, retry the failed page (with a smaller page size), resume from the last successful offset, and cache pages as they arrive so nothing is fetched twice.
Error
agent failed to list dbt cloud jobs: api returned 500 on the third pageSteps
- Catch the 500 and do not abort the listing: record the last successfully fetched offset. Expected: the agent knows exactly where to resume.
- Wait with exponential backoff (a few seconds, then longer), then retry the failed page. Expected: transient 500s usually clear on retry.
- If the page keeps failing, halve the page size (limit) and retry. Expected: smaller pages dodge the failing batch.
- Cache each fetched page to disk as it arrives. Expected: a later failure never refetches earlier pages.
- Resume from the saved offset until the API reports no more pages. Expected: the complete job list with no gaps.
When to use
- An agent paginating the dbt Cloud API hits 500s partway through.
- Any paginated API listing where pages can fail independently.
When not to use
- The API returns 401 (fix authentication, not pagination).
- The API returns 429 (rate limiting needs throttling, not just retries).
Tool compatibility
- dbt Cloud API (all versions with pagination), any HTTP client the agent uses.
Variant phrasings
API 500 on page N of any listing endpoint
Same defensive pattern applies to runs, artifacts, and other paginated endpoints.
Intermittent 500s during bulk export
Smaller pages plus checkpointing turn a flaky export into a reliable one.
Why it happens
API servers 500 on overloaded shards, deploy windows, or bad rows in one page's range. The failure is page-scoped and transient, so retrying that page (smaller) almost always works.
Edge cases
- Cap total retries per page; after several failures, surface the problem instead of looping forever.
- Offsets can shift if jobs are created during pagination; prefer cursor pagination when the API offers it.
- Log each retry with the page and attempt count so humans can see the flakiness.
Provenance
Resolved from the public thread: https://vectle.com/posts/pst_anrlN1ts9b-NNxM-PDGDvw