agent hallucinated aria fix that breaks focus order
Fixes remediation agents landing harmful aria patches by gating every patch on an axe re-scan. Use it when agent-proposed fixes introduce new violations. Not for human-authored fixes, which need review but not this gate.
agent hallucinated aria fix that breaks focus order - how to fix it
TL;DR
Verify every agent-proposed aria fix with a re-scan before accepting it: run axe on the patched DOM and diff the violation list, rejecting any patch that trades one violation for another. Agents confidently invent aria that looks right and breaks focus. One line of why: ARIA is a contract with assistive tech, and a plausible-sounding attribute in the wrong place silently rewires the page.
The error, verbatim
RemediationAgentWarning: proposed fix introduced new violations
patch: added tabindex=1 to nav landmarks
new_violations: focus-order-semantics, tabindex-no-positive (lint)
original_violation: region (fixed, but regressed)
Fix it step by step
Step 1: Reproduce the regression
node agent/remediate.js --dry-run | rg -i 'hallucinat|regress|new violation' | head -5Expected: The agent proposes a fix that introduces new violations.
Step 2: Re-scan the patched DOM
npx @axe-core/cli https://example.com/ --save axe-after.json | tail -2Expected: Confirms the patch trades violations instead of clearing them.
Step 3: Add the verify-before-accept gate
rg -n 'verify|re-scan|accept' agent/remediate.js | head -10Expected: Add: apply patch in a sandbox DOM, re-run axe, accept only if total violations strictly decrease.
Step 4: Re-run remediation
node agent/remediate.js | tail -4Expected: Agent proposes, verifies, and only lands patches that reduce violations.
Step 5: Add a regression probe
node agent/run-audit.js --smoke | tail -3Expected: Smoke run passes; schedule it so the breakdown is caught if it ever regresses.
When to use this skill
- You run an agent that scans UIs for accessibility and it hits this breakdown
- The agent's scan loop stalls, crashes, or loops on this exact failure
- You are hardening an audit agent's error handling for production scans
When NOT to use this skill
- A human runs the scan manually and it works, this is agent-harness failure handling
- The scan completes and only reports violations, use the rule-specific skills
Compatibility
Remediation agent with axe-core re-scan as the acceptance gate. The sandbox can be a headless page or JSDOM plus axe. Pin the tool version in the lockfile so scans stay reproducible across machines.
Variant phrasings
agent aria fix regression
Same pattern, any invented aria that backfires.
remediation made accessibility worse
Practitioner phrasing for the verify gate.
the breakdown hits other routes too
Agent failure modes are systemic; apply the hardening to every route the agent covers, not just the one that failed.
Why it happens
Language models know ARIA vocabulary but not ARIA semantics, so they propose fixes like tabindex=1 on landmarks or role=button on divs that read plausibly and fail rules. Without a verification gate, the agent lands the patch and reports success. The gate is simple and non-negotiable: every patch gets re-scanned, and acceptance requires the violation count to strictly decrease with no new critical or serious violations. Anything else is rejected with the diff as feedback for the next proposal. Agent breakdowns are systemic: the same failure mode will hit every route, page, or run the agent touches. Harden the harness once (timeouts, loop detection, verification gates) instead of patching per page, and keep breakdown telemetry separate from violation counts.
Edge cases
- Never let the agent edit aria and skip verification because the change looks small, small aria changes have big effects.
- Keep a denylist of known-bad patterns (positive tabindex, nested roles) as a fast pre-check before the full re-scan.
- The verify gate needs the same axe version and rules as the original scan, or diffs are meaningless.
- Log breakdowns separately from violations in agent telemetry; mixing them hides whether the agent itself is getting more reliable.
Provenance
Resolved from the public thread: https://vectle.com/posts/pstZU72Q-bBhHOCqB9C-d8_g