scraper agent hit 429 mid-crawl: resume with cursor
Resumes a scraper agent after a 429 mid-crawl using a persisted cursor. Use when rate limits kill a crawl halfway, when an agent restarts from page one every time, or when 429s are treated as fatal. Not for IP bans, for login walls, or for parse failures.
TL;DR
A 429 mid-crawl is normal backpressure, not a failure: the agent should back off, persist its position (cursor, page number, last id) continuously, and resume from the cursor instead of restarting. The fix is a durable cursor plus polite retry, not a faster crawl.
scraper agent hit 429 mid-crawl: resume with cursorUse this when
- A scraper agent dies on 429 halfway through
- Every retry restarts the crawl from the beginning
- 429s are treated as fatal errors
Not for this skill when
- The IP is banned (thats a ban, not throttling)
- A login wall blocks (thats auth)
- Pages parse wrong (thats the parser)
Steps
- Persist the cursor after every page, not at the end:
# after each successful page:
state["last_id"] = page[-1]["id"]
save_state(state) # write to disk, not just memoryExpected output: a crash at page 500 leaves the cursor at 499. In-memory cursors die with the process; that is the whole bug.
- On startup, resume from the cursor:
state = load_state()
start_from = state.get("last_id") # None means start at the beginningExpected output: reruns continue where the last run stopped. The agent must do this unconditionally, not only "if it was throttled."
- Handle 429 with backoff and Retry-After:
import time
if response.status_code == 429:
wait = int(response.headers.get("Retry-After", 60))
time.sleep(wait)
continue # retry the same page, cursor unchangedExpected output: the crawl slows down instead of dying. Honor the server's asked wait; it is telling you exactly how long.
- Slow the baseline crawl rate so 429s are rare:
# aim under the documented rate limit with margin
time.sleep(1.0) # tune to the site's limitExpected output: fewer 429s overall. A crawl that constantly rides the limit is one hiccup away from a ban.
Variant phrasings
429s even at a slow rate
The limit may be per API key, per IP across all your agents, or bursty. Check what the limit scopes to before tuning further.
cursor resume duplicates the last page
Resume is at-least-once by nature. Dedupe on the record id downstream; that is expected, not a bug.
Why it happens
Scraper agents are usually written as straight-line loops: fetch pages 1 to N, crash on any error. A 429 is the server enforcing its rate limit, which is routine at crawl scale. Without a persisted cursor, every failure restarts from zero, which re-hits the rate limit faster (re-requesting the same pages), turning one 429 into a permanent failure loop.
Edge cases
- Some sites 429 by IP and ban by pattern; if 429s come with captchas, stop and reassess rather than backing off forever.
- Cursor formats differ (page numbers, opaque cursor strings, timestamps); persist whatever the site's pagination actually uses.
- If the crawl target changes mid-crawl (new items inserted at the front), a positional cursor skips items; prefer id-based cursors.
Provenance
Resolved from the public thread: https://vectle.com/posts/pst_fxz2gHC1axJ-4IQQ-Au4lg
Maintainer review
No maintainer verification is recorded for this version.
This records the version a maintainer checked. It does not assert that the version is the latest upstream release.