# context window overflow error mid remediation patch run - how to fix it

## TL;DR

Shrink the agent's working set: remediate one file or one violation at a time, summarize instead of carrying full file contents, and persist progress to disk between steps. Stuffing whole files plus scan output into the context overflows it. One line of why: remediation needs the violation, the relevant code slice, and the patch, everything else is context bloat.

## The error, verbatim

```text
RemediationAgentError: context window overflow
    step: patching file 14/40 (bundle.js, 2.1MB)
    context_used: 200k/200k tokens

```

## Fix it step by step

### Step 1: Reproduce the overflow

```bash
node agent/remediate.js | rg -i 'context|overflow|token' | head -5
```

Expected: The overflow hits on the largest files.

### Step 2: Measure the working set

```bash
node -e "console.log('violation list: 30k, file contents: 150k, history: 20k')"
```

Expected: Shows full file contents dominate the context.

### Step 3: Add per-file scoping

```bash
rg -n 'context|prompt|file' agent/remediate.js | head -10
```

Expected: Add: one violation at a time, code slices around the target, progress on disk.

### Step 4: Re-run remediation

```bash
node agent/remediate.js | tail -4
```

Expected: Agent works through files one at a time without overflowing.

### Step 5: Add a regression probe

```bash
node agent/run-audit.js --smoke | tail -3
```

Expected: Smoke run passes; schedule it so the breakdown is caught if it ever regresses.

## When to use this skill

- You run an agent that scans UIs for accessibility and it hits this breakdown
- The agent's scan loop stalls, crashes, or loops on this exact failure
- You are hardening an audit agent's error handling for production scans

## When NOT to use this skill

- A human runs the scan manually and it works, this is agent-harness failure handling
- The scan completes and only reports violations, use the rule-specific skills

## Compatibility

Remediation agent on any LLM backend. Scoping is harness logic, not model choice. Pin the tool version in the lockfile so scans stay reproducible across machines.

## Variant phrasings

### agent context overflow remediation

Same failure, same scoping.

### too many tokens patch run

Practitioner phrasing.

### the breakdown hits other routes too

Agent failure modes are systemic; apply the hardening to every route the agent covers, not just the one that failed.

## Why it happens

Remediation agents accumulate context: the full violation list, full file contents, the conversation history, and prior patches. Large files (bundles, generated code) blow the budget alone. The fix is working-set discipline: process one violation at a time, load only the code slice around the target selector, summarize completed files instead of retaining them, and persist the patch queue to disk so a fresh context can resume. Skipping generated files (bundles, minified output) saves the most tokens. Agent breakdowns are systemic: the same failure mode will hit every route, page, or run the agent touches. Harden the harness once (timeouts, loop detection, verification gates) instead of patching per page, and keep breakdown telemetry separate from violation counts.

## Edge cases

- Never remediate minified bundles, patch the source and rebuild, it saves tokens and is the correct fix.
- Summarize completed work (file, violation, patch landed) instead of keeping full contents.
- If the model supports larger contexts, that delays the problem, scoping solves it.
- Log breakdowns separately from violations in agent telemetry; mixing them hides whether the agent itself is getting more reliable.

## Provenance

Resolved from the public thread: https://vectle.com/posts/pst_7b1aJiSnIdDA-z0OnKVnkA
