design api rate limit hit mid alt-text generation run failed
Fixes remediation agent 429s during alt-text generation with backoff, caching, and batching. Use it when the design API limits the agent mid-gallery. Not for model-side token limits, which are context problems.
design api rate limit hit mid alt-text generation run failed - how to fix it
TL;DR
Treat the design API like any rate-limited dependency: back off with jitter, cache aggressively, and batch requests. Alt-text generation that calls a design API per image will hit the limit on any real gallery. One line of why: the limit is per minute, the gallery is hundreds of images, and the math never works without batching and caching.
The error, verbatim
RemediationAgentError: design API rate limit hit (429)
generated: 40/480 alt texts, cache: none
retry-after: 60, backoff: none (aborted)
Fix it step by step
Step 1: Reproduce the 429
node agent/remediate.js --alt-text | rg -i '429|rate' | head -5Expected: The limit hits partway through the gallery.
Step 2: Check the API quota
node -e "console.log('quota: 60/min, images: 480, calls per image: 1, time needed without batching: 8min+')"Expected: The math shows batching is required.
Step 3: Add backoff, cache, and batching
rg -n '429|cache|batch' agent/remediate.js | head -10Expected: Add: honor retry-after, cache by image hash, batch where the API allows.
Step 4: Re-run remediation
node agent/remediate.js --alt-text | tail -4Expected: Agent paces itself through the gallery and completes.
Step 5: Add a regression probe
node agent/run-audit.js --smoke | tail -3Expected: Smoke run passes; schedule it so the breakdown is caught if it ever regresses.
When to use this skill
- You run an agent that scans UIs for accessibility and it hits this breakdown
- The agent's scan loop stalls, crashes, or loops on this exact failure
- You are hardening an audit agent's error handling for production scans
When NOT to use this skill
- A human runs the scan manually and it works, this is agent-harness failure handling
- The scan completes and only reports violations, use the rule-specific skills
Compatibility
Remediation agent calling a design or vision API. Cache key is the image content hash. Pin the tool version in the lockfile so scans stay reproducible across machines.
Variant phrasings
agent 429 alt text api
Same limit, same pacing.
design api quota alt generation
Practitioner phrasing.
the breakdown hits other routes too
Agent failure modes are systemic; apply the hardening to every route the agent covers, not just the one that failed.
Why it happens
Alt-text generation leans on design or vision APIs with per-minute quotas, and galleries need hundreds of calls. Agents without backoff abort on the first 429 and lose progress; agents without caching re-request the same images every run. The fix is the standard trio: honor retry-after with exponential backoff and jitter, cache results by image hash across runs, and batch requests where the API supports it. Pace the run under the quota instead of bouncing off it. Agent breakdowns are systemic: the same failure mode will hit every route, page, or run the agent touches. Harden the harness once (timeouts, loop detection, verification gates) instead of patching per page, and keep breakdown telemetry separate from violation counts.
Edge cases
- Cache by content hash, not URL, galleries reuse images under different URLs.
- Jitter backoff so parallel agent workers do not retry in lockstep.
- If the quota is daily and the gallery is huge, split the run across days with checkpoints.
- Log breakdowns separately from violations in agent telemetry; mixing them hides whether the agent itself is getting more reliable.
Provenance
Resolved from the public thread: https://vectle.com/posts/pst6hP7Y0NwHTnOgWM5hsQhA