# axe api rate limit hit mid full-site audit run failed - how to fix it

## TL;DR

Back off and resume: when the audit API rate-limits the agent mid-crawl, the agent must pause with exponential backoff, persist the crawl frontier, and resume where it stopped instead of restarting. One line of why: a full-site audit that restarts from zero on every 429 never finishes, and hammering the limiter gets the agent blocked.

## The error, verbatim

```text
AuditAgentError: API rate limit hit mid-audit (429)
    crawled: 340/1200 urls, frontier lost on abort
    retry-after: 60

```

## Fix it step by step

### Step 1: Reproduce the 429

```bash
node agent/crawl.js --limit 50 | rg -i '429|rate' | head -5
```

Expected: The rate limit triggers under crawl load.

### Step 2: Persist the frontier

```bash
rg -n 'frontier|queue|visited' agent/crawl.js | head -10
```

Expected: Add frontier persistence (visited set plus queue) to disk after each page.

### Step 3: Add backoff and resume

```bash
node agent/crawl.js --resume | tail -5
```

Expected: Agent honors retry-after, backs off exponentially, and resumes from the saved frontier.

### Step 4: Complete the audit

```bash
node agent/crawl.js --resume | rg -i 'complete|crawled' | tail -3
```

Expected: Crawl finishes all 1200 URLs across multiple rate-limit windows.

### Step 5: Add a regression probe

```bash
node agent/run-audit.js --smoke | tail -3
```

Expected: Smoke run passes; schedule it so the breakdown is caught if it ever regresses.

## When to use this skill

- You run an agent that scans UIs for accessibility and it hits this breakdown
- The agent's scan loop stalls, crashes, or loops on this exact failure
- You are hardening an audit agent's error handling for production scans

## When NOT to use this skill

- A human runs the scan manually and it works, this is agent-harness failure handling
- The scan completes and only reports violations, use the rule-specific skills

## Compatibility

Audit agent with a crawl frontier. Backoff honors the retry-after header; persistence is a JSON file on disk. Pin the tool version in the lockfile so scans stay reproducible across machines.

## Variant phrasings

### agent 429 mid audit

Same failure, same backoff-and-resume.

### crawl rate limited audit

Practitioner phrasing.

### the breakdown hits other routes too

Agent failure modes are systemic; apply the hardening to every route the agent covers, not just the one that failed.

## Why it happens

Full-site audits crawl hundreds of URLs, and audit APIs rate-limit aggressively. Agents without frontier persistence lose all progress on a 429 and restart from zero, which re-triggers the limiter faster each time, a death spiral. The fix has three parts: persist the visited set and queue continuously, honor retry-after with exponential backoff and jitter, and resume from disk. Also lower the crawl concurrency to stay under the limit instead of bouncing off it. Agent breakdowns are systemic: the same failure mode will hit every route, page, or run the agent touches. Harden the harness once (timeouts, loop detection, verification gates) instead of patching per page, and keep breakdown telemetry separate from violation counts.

## Edge cases

- Jitter the backoff so multiple agent instances do not retry in lockstep.
- Lower concurrency proactively (e.g. 2 requests per second) instead of relying on 429s as the signal.
- Some APIs rate-limit per key per minute and per day, track both budgets in the agent.
- Log breakdowns separately from violations in agent telemetry; mixing them hides whether the agent itself is getting more reliable.

## Provenance

Resolved from the public thread: https://vectle.com/posts/pst_6VKyPbeN04yM8CMBpegZlQ
