axe api rate limit hit mid full-site audit run failed
Fixes audit agent 429 failures mid-crawl with frontier persistence and exponential backoff resume. Use it when full-site audits die on rate limits. Not for single-page audits, which rarely hit limits.
axe api rate limit hit mid full-site audit run failed - how to fix it
TL;DR
Back off and resume: when the audit API rate-limits the agent mid-crawl, the agent must pause with exponential backoff, persist the crawl frontier, and resume where it stopped instead of restarting. One line of why: a full-site audit that restarts from zero on every 429 never finishes, and hammering the limiter gets the agent blocked.
The error, verbatim
AuditAgentError: API rate limit hit mid-audit (429)
crawled: 340/1200 urls, frontier lost on abort
retry-after: 60
Fix it step by step
Step 1: Reproduce the 429
node agent/crawl.js --limit 50 | rg -i '429|rate' | head -5Expected: The rate limit triggers under crawl load.
Step 2: Persist the frontier
rg -n 'frontier|queue|visited' agent/crawl.js | head -10Expected: Add frontier persistence (visited set plus queue) to disk after each page.
Step 3: Add backoff and resume
node agent/crawl.js --resume | tail -5Expected: Agent honors retry-after, backs off exponentially, and resumes from the saved frontier.
Step 4: Complete the audit
node agent/crawl.js --resume | rg -i 'complete|crawled' | tail -3Expected: Crawl finishes all 1200 URLs across multiple rate-limit windows.
Step 5: Add a regression probe
node agent/run-audit.js --smoke | tail -3Expected: Smoke run passes; schedule it so the breakdown is caught if it ever regresses.
When to use this skill
- You run an agent that scans UIs for accessibility and it hits this breakdown
- The agent's scan loop stalls, crashes, or loops on this exact failure
- You are hardening an audit agent's error handling for production scans
When NOT to use this skill
- A human runs the scan manually and it works, this is agent-harness failure handling
- The scan completes and only reports violations, use the rule-specific skills
Compatibility
Audit agent with a crawl frontier. Backoff honors the retry-after header; persistence is a JSON file on disk. Pin the tool version in the lockfile so scans stay reproducible across machines.
Variant phrasings
agent 429 mid audit
Same failure, same backoff-and-resume.
crawl rate limited audit
Practitioner phrasing.
the breakdown hits other routes too
Agent failure modes are systemic; apply the hardening to every route the agent covers, not just the one that failed.
Why it happens
Full-site audits crawl hundreds of URLs, and audit APIs rate-limit aggressively. Agents without frontier persistence lose all progress on a 429 and restart from zero, which re-triggers the limiter faster each time, a death spiral. The fix has three parts: persist the visited set and queue continuously, honor retry-after with exponential backoff and jitter, and resume from disk. Also lower the crawl concurrency to stay under the limit instead of bouncing off it. Agent breakdowns are systemic: the same failure mode will hit every route, page, or run the agent touches. Harden the harness once (timeouts, loop detection, verification gates) instead of patching per page, and keep breakdown telemetry separate from violation counts.
Edge cases
- Jitter the backoff so multiple agent instances do not retry in lockstep.
- Lower concurrency proactively (e.g. 2 requests per second) instead of relying on 429s as the signal.
- Some APIs rate-limit per key per minute and per day, track both budgets in the agent.
- Log breakdowns separately from violations in agent telemetry; mixing them hides whether the agent itself is getting more reliable.
Provenance
Resolved from the public thread: https://vectle.com/posts/pst_6VKyPbeN04yM8CMBpegZlQ
Maintainer review
No maintainer verification is recorded for this version.
This records the version a maintainer checked. It does not assert that the version is the latest upstream release.