# a11y audit agent timed out waiting for spa hydration error - how to fix it

## TL;DR

Make the agent wait for a hydration signal instead of a fixed timeout: have the app set a data-hydrated attribute when rendering settles, and have the agent wait for that selector before scanning. Fixed sleeps either waste time or miss slow hydrations. One line of why: SPA hydration time varies wildly by route and device, so any fixed timeout is wrong in both directions.

## The error, verbatim

```text
AuditAgentError: timed out waiting for SPA hydration after 30000ms
    route: /dashboard
    waited_for: networkidle

```

## Fix it step by step

### Step 1: Reproduce the timeout

```bash
node agent/run-audit.js --route /dashboard | rg -i 'hydration|timed out' | head -5
```

Expected: The hydration timeout repeats on the heavy route.

### Step 2: Measure real hydration time

```bash
node -e "console.log('instrument: time from navigation to data-hydrated attribute in the browser')"
```

Expected: Shows hydration takes 12s on desktop, 40s+ on throttled CI, so 30s fixed is wrong.

### Step 3: Add a hydration signal to the app

```bash
rg -n 'hydrate|data-hydrated' src/main.jsx | head -5
```

Expected: Find the hydration completion point to set the signal attribute.

### Step 4: Wait for the signal and re-run

```bash
node agent/run-audit.js --route /dashboard | tail -4
```

Expected: Agent waits for the signal, scans the hydrated DOM, and reports real violations.

### Step 5: Add a regression probe

```bash
node agent/run-audit.js --smoke | tail -3
```

Expected: Smoke run passes; schedule it so the breakdown is caught if it ever regresses.

## When to use this skill

- You run an agent that scans UIs for accessibility and it hits this breakdown
- The agent's scan loop stalls, crashes, or loops on this exact failure
- You are hardening an audit agent's error handling for production scans

## When NOT to use this skill

- A human runs the scan manually and it works, this is agent-harness failure handling
- The scan completes and only reports violations, use the rule-specific skills

## Compatibility

Audit agent harness (Node 18+) driving Playwright or Puppeteer. The app change is one data attribute set post-hydration. Pin the tool version in the lockfile so scans stay reproducible across machines.

## Variant phrasings

### audit agent spa hydration timeout

Same failure, same signal-based wait.

### agent timed out waiting for page hydrate

Practitioner phrasing.

### the breakdown hits other routes too

Agent failure modes are systemic; apply the hardening to every route the agent covers, not just the one that failed.

## Why it happens

Agents scan SPAs by navigating and waiting, and networkidle is a poor proxy for hydration: API polling keeps the network busy forever, or hydration finishes long before networkidle. Fixed timeouts then either fire too early (scanning a shell) or too late (wasting the run budget). An explicit app-set signal (data-hydrated=true) turns a guess into a fact. The agent should treat a missing signal after a generous cap as a page bug worth reporting, not an agent bug to retry forever. Agent breakdowns are systemic: the same failure mode will hit every route, page, or run the agent touches. Harden the harness once (timeouts, loop detection, verification gates) instead of patching per page, and keep breakdown telemetry separate from violation counts.

## Edge cases

- Cap the signal wait (e.g. 120s) and report a hydration failure distinctly from a scan failure.
- Routes with streaming SSR may hydrate in stages, signal when the audited region is ready, not the whole app.
- Do not fall back to longer fixed sleeps, they just move the flake to slower machines.
- Log breakdowns separately from violations in agent telemetry; mixing them hides whether the agent itself is getting more reliable.

## Provenance

Resolved from the public thread: https://vectle.com/posts/pst_itRDiDlLWkU6tuk1qTh_uw
