# docs agent failed after context window overflow mid api reference generation

## TL;DR
Generate the reference one module at a time and merge the parts instead of feeding the whole codebase in one go. The agent stuffed every source file into a single pass and ran out of room halfway, leaving a truncated half-written reference. Chunked generation with a per-module budget finishes reliably.

## The error

```text
docs agent failed after context window overflow mid api reference generation
```

## Steps

1. Measure the input. Count the source files and total size feeding the generation step and compare against the model's working budget.

Expected: you confirm the input exceeds the budget, which explains the mid-run truncation.

2. Split generation by module. Make one agent call per top-level module or package, each producing only that module's reference section.

Expected: each call's input fits comfortably with room to spare for the output.

3. Write each module's output to its own file, then concatenate the parts behind a generated index page.

Expected: the full reference exists as the union of the parts, with no truncated pages.

4. Cap per-module input explicitly. Skip files over a size threshold with a short summarized note, and set a max output size per call.

Expected: no single call can overflow again, even as the repo grows.

## Use this when

- reference generation dies partway or emits truncated pages on large codebases
- the output stops mid-module with no error, just cut off
- adding more source files makes previously working generation fail

## Not for this skill when

- the overflow happens in a different pipeline stage, like changelog or summary generation
- the repo is small enough that one pass fits fine (look for a different cause)
- generation fails with an error rather than truncation

## Variant phrasings

### api docs generation ran out of context
Split by module and merge; the single-pass approach does not scale.

### reference writer truncated on big repo
Check which module it died on; that module probably needs splitting too.

### llm docs output cut off halfway
Measure input size per call and cap it; truncation without errors is the signature of overflow.

## Why it happens
Reference generation is the most input-hungry step in the pipeline: it wants every source file at once to cross-link correctly. Whole-repo single-pass prompts grow with codebase size until they exceed the model's budget, and the run dies or truncates wherever the limit lands.

## Edge cases

- Cross-module links need stable anchors. Generate the anchor map first, then write each module against it, or links between parts will break.
- Skipped large files still need a stub entry so the reference does not silently omit public API.
- Merging must use an explicit module order or diffs churn every run. Sort modules deterministically.
- Very deep module trees can overflow even per-module; split large modules by file as a second level of chunking.

## Provenance

Resolved from the public thread: https://vectle.com/posts/pst_H7St3WZquOERYu8VXZblfw
