# Filing agent broke on scanned PDF image pages

## TL;DR
Scanned image pages break filing agents because the parser expects text and gets pictures. The fix is a page-type check before parsing: route text pages to the text extractor and image pages to OCR. Mixed filings need per-page routing, and the agent should never assume a filing is one or the other.

## The error
```text
(parse failed)
filing agent broke on scanned pdf image pages; text extractor returned empty, downstream parse failed
```

## When this helps
- a filing agent breaks on scanned pages
- mixed text and image PDFs fail parsing
- building per-page extraction routing
- validating filing text coverage

## When it doesn't
- the whole filing is scanned; OCR the whole thing, no routing needed
- OCR output is garbage; the scan quality is the problem, find a better copy
- you need the images themselves; extraction gives text, not figures

## Works with
python 3.8+ with pymupdf; ocrmypdf with tesseract.

## Steps
### 1. Classify each page as text or image before parsing
```python
import fitz
doc = fitz.open("filing.pdf")
for i, page in enumerate(doc[:5]):
    text = page.get_text().strip()
    imgs = len(page.get_images())
    kind = "text" if len(text) - 100 in range(1, 10**9) else ("image" if imgs else "empty")
    print("page", i + 1, kind)
```
Expected: A per-page classification. The agent routes each page to the right extractor instead of guessing once.

### 2. OCR the image pages
```bash
ocrmypdf --force-ocr filing.pdf filing_ocr.pdf
python3 -c "import fitz; d=fitz.open('filing_ocr.pdf'); print('ocred pages:', len(d), 'page1 chars:', len(d[0].get_text()))"
```
Expected: An OCRed filing. Force-ocr handles mixed filings by adding a text layer where it is missing.

### 3. Parse text pages and OCRed pages through one path
```python
import fitz
doc = fitz.open("filing_ocr.pdf")
texts = [p.get_text() for p in doc]
print("pages:", len(texts), "total chars:", sum(len(t) for t in texts))
```
Expected: Uniform text output. After OCR, every page parses through the same code path.

### 4. Validate page coverage before downstream parsing
```python
import fitz
doc = fitz.open("filing_ocr.pdf")
empty = [i for i, p in enumerate(doc) if len(p.get_text().strip()) == 0]
print("empty pages:", empty[:10])
print("flagged for manual review" if empty else "full coverage")
```
Expected: A coverage report. Pages still empty after OCR are diagrams or hopeless scans; the briefing notes them.

## Other ways people phrase this
### scanned pdf pages parse failed agent
Classify per page, OCR the image ones, parse uniformly after.

### filing agent empty text image pages
The extractor is fine; the pages have no text. OCR first.

### mixed pdf text and scanned pages
Per-page routing. Never assume the whole filing is one type.

## Why it happens
Filings mix born-digital pages with scanned exhibits, and text extractors return empty for the scanned ones. Agents that extract the whole file in one call get partial text and fail downstream. Per-page classification plus OCR normalizes the filing before any parsing happens.

## Edge cases
- OCR on already-text pages can degrade quality; force-ocr only where needed.
- Diagrams and charts stay empty after OCR; note them rather than failing.
- Some scanned pages are sideways; detect orientation before OCR.
- Validate a known figure after OCR; numbers are what OCR mangles first.

## Provenance

Resolved from the public thread: https://vectle.com/posts/pst_JGHb1c6EdK1PAnsELuv9Ow
