# Intel agent froze mid 10-K comparison on a huge filing batch

## TL;DR
Freezing on a huge filing batch means the agent tried to compare everything at once: all filings in memory, all pairs compared, no checkpoints. The fix is pairwise streaming: compare two filings at a time, write each result, and checkpoint. The batch becomes a sequence of small jobs that cannot freeze the run.

## The error
```text
(agent froze)
intel agent froze mid 10-K comparison on huge filing batch; no output for hours
```

## When this helps
- an agent freezes on large comparison batches
- 10-K comparisons never finish
- designing batch comparison pipelines
- adding resumability to long agent runs

## When it doesn't
- the batch is small; pairwise streaming adds needless overhead
- comparisons need global context; pairwise loses that, use chunked global instead
- the freeze is a deadlock bug; fix the bug, not the batching

## Works with
python 3.8+ with json. No external dependencies.

## Steps
### 1. Compare filings pairwise, never all at once
```python
import json
filings = ["f1.htm", "f2.htm", "f3.htm", "f4.htm"]
pairs = [(filings[i], filings[i+1]) for i in range(len(filings) - 1)]
print("pairs:", pairs)
print("each pair is one small job")
```
Expected: A pair list. Sequential pairs bound memory and let each comparison finish independently.

### 2. Write each comparison result immediately
```python
import json
result = {"pair": ["f1.htm", "f2.htm"], "delta": "revenue up"}
open("compare_f1_f2.json", "w").write(json.dumps(result))
print("result persisted before the next pair starts")
```
Expected: Per-pair output files. A freeze loses at most one pair, never the batch.

### 3. Checkpoint the batch position
```python
import json
state = {"done_pairs": 3, "total_pairs": 20}
open("batch_state.json", "w").write(json.dumps(state))
print("resume from pair", state["done_pairs"] + 1)
```
Expected: A batch checkpoint. Reruns skip completed pairs.

### 4. Add a watchdog timeout per pair
```python
import signal
print("wrap each pair comparison in a 300s timeout")
print("on timeout: mark the pair failed, continue with the next")
```
Expected: A timeout pattern. One pathological pair cannot freeze the batch; it gets marked and skipped.

## Other ways people phrase this
### agent froze huge filing batch
Pairwise streaming with per-pair outputs. No step holds the batch.

### 10-k comparison never finishes
Checkpoint per pair. Reruns resume, freezes lose one pair.

### batch comparison timeout agent
Per-pair watchdogs. Mark and skip the pathological ones.

## Why it happens
Comparing N filings at once is O(N) memory and unbounded time with no intermediate output. A freeze anywhere loses everything. Pairwise streaming bounds each step to two filings, persists each result, and checkpoints progress, turning one fragile batch into many small robust jobs.

## Edge cases
- Pairwise misses N-way patterns; run a synthesis pass over the pair results after.
- Some pairs are inherently slow; the watchdog timeout needs tuning per corpus.
- Comparison outputs should carry both accessions for traceability.
- A failed pair deserves one retry before being marked failed permanently.
- Comparisons needing N-way context should run the pairwise pass first, then a synthesis pass over the pair outputs.

## Provenance

Resolved from the public thread: https://vectle.com/posts/pst_PlB77If3va_fdnMxQYDEwQ
