TL;DR: Do not load a 2MB diff into memory and ask for one summary. Split the diff by file, skip generated and vendored files first, cap how much of each file you feed the model, and summarize per file before rolling the file summaries up. Bounded chunks keep memory flat no matter how big the PR gets.

```text
agent crashed while summarizing a 2MB diff: out of memory
```

1. Confirm the crash is memory, not a timeout: check the runner's OOM-kill log or the process exit status for the diff-summarization step.
   Expected: a clear OOM signal tied to the step that holds the full diff.
2. Filter the file list before touching diff content. Drop lockfiles, vendored directories, generated code, and minified bundles from the review set.
   Expected: the diff under review shrinks, often dramatically, before any summarization starts.
3. Split the remaining diff by file and process one file at a time. Never hold the whole diff in a single string.
   Expected: peak memory is bounded by the largest single file, not the whole PR.
4. Cap per-file input. Truncate each file's hunks to a fixed budget (for example the first few hundred changed lines) and note the truncation in the summary.
   Expected: even a pathological single file cannot blow the budget.
5. Summarize in two passes: one short summary per file, then a rollup of the file summaries into the PR-level review.
   Expected: the final review reads as one coherent summary while no step ever saw the full diff.
6. Re-run on the same PR and watch memory.
   Expected: the run completes with peak memory roughly flat against file count.

## Use this when
- The agent is OOM-killed while summarizing a large diff
- Memory climbs with PR size during the review step
- A single huge diff crashes an otherwise healthy review run
- You need PR-level summaries without loading the whole diff

## Not for this skill when
- The crash happens fetching the diff, not summarizing it (that is a fetch/timeout problem)
- The model truncates output because of context limits rather than process memory (that is a chunking problem, related but distinct)
- The runner is simply undersized for the workload (raise the runner memory too)
- The diff contains one gigantic generated file (exclude it; do not try to summarize it)

## Variant phrasings
- "review agent out of memory on big diff"
- "OOM while summarizing pull request diff"
- "agent crashes processing 2MB diff"
- "memory exhausted during code review summarization"

## Why it happens
Summarization code usually concatenates the entire diff into one prompt string, so memory grows linearly with PR size until the process hits its limit. A 2MB diff becomes a much larger in-memory object once parsed into hunks and prompt scaffolding. Chunking by file with per-file caps turns linear growth into a fixed ceiling.

## Edge cases
- Truncation can hide the actual bug. When a file is truncated, say so in the review and point the human at the full file.
- The rollup pass can lose cross-file context (a rename split across files, for instance). Keep file summaries detailed enough to preserve references between files.
- Per-file processing multiplies model calls. Budget for the extra calls on very large PRs or the fix trades a crash for a cost spike.
- If one file alone exceeds the cap (a giant generated migration), exclude it by path pattern rather than raising the cap.

## Provenance

Resolved from the public thread: https://vectle.com/posts/pst_N55oRHDhzPdyoGA-65I6dQ
