filing agent broke on scanned pdf image pages, parse failed
This skill fixes filing agents that break on scanned PDF image pages. Use it when mixed filings fail parsing or when building extraction routing. It is not for fully scanned filings; the fix is per-page text-or-image classification, OCR for image pages, and coverage validation.
Filing agent broke on scanned PDF image pages
TL;DR
Scanned image pages break filing agents because the parser expects text and gets pictures. The fix is a page-type check before parsing: route text pages to the text extractor and image pages to OCR. Mixed filings need per-page routing, and the agent should never assume a filing is one or the other.
The error
(parse failed)
filing agent broke on scanned pdf image pages; text extractor returned empty, downstream parse failedWhen this helps
- a filing agent breaks on scanned pages
- mixed text and image PDFs fail parsing
- building per-page extraction routing
- validating filing text coverage
When it doesn't
- the whole filing is scanned; OCR the whole thing, no routing needed
- OCR output is garbage; the scan quality is the problem, find a better copy
- you need the images themselves; extraction gives text, not figures
Works with
python 3.8+ with pymupdf; ocrmypdf with tesseract.
Steps
1. Classify each page as text or image before parsing
import fitz
doc = fitz.open("filing.pdf")
for i, page in enumerate(doc[:5]):
text = page.get_text().strip()
imgs = len(page.get_images())
kind = "text" if len(text) - 100 in range(1, 10**9) else ("image" if imgs else "empty")
print("page", i + 1, kind)Expected: A per-page classification. The agent routes each page to the right extractor instead of guessing once.
2. OCR the image pages
ocrmypdf --force-ocr filing.pdf filing_ocr.pdf
python3 -c "import fitz; d=fitz.open('filing_ocr.pdf'); print('ocred pages:', len(d), 'page1 chars:', len(d[0].get_text()))"Expected: An OCRed filing. Force-ocr handles mixed filings by adding a text layer where it is missing.
3. Parse text pages and OCRed pages through one path
import fitz
doc = fitz.open("filing_ocr.pdf")
texts = [p.get_text() for p in doc]
print("pages:", len(texts), "total chars:", sum(len(t) for t in texts))Expected: Uniform text output. After OCR, every page parses through the same code path.
4. Validate page coverage before downstream parsing
import fitz
doc = fitz.open("filing_ocr.pdf")
empty = [i for i, p in enumerate(doc) if len(p.get_text().strip()) == 0]
print("empty pages:", empty[:10])
print("flagged for manual review" if empty else "full coverage")Expected: A coverage report. Pages still empty after OCR are diagrams or hopeless scans; the briefing notes them.
Other ways people phrase this
scanned pdf pages parse failed agent
Classify per page, OCR the image ones, parse uniformly after.
filing agent empty text image pages
The extractor is fine; the pages have no text. OCR first.
mixed pdf text and scanned pages
Per-page routing. Never assume the whole filing is one type.
Why it happens
Filings mix born-digital pages with scanned exhibits, and text extractors return empty for the scanned ones. Agents that extract the whole file in one call get partial text and fail downstream. Per-page classification plus OCR normalizes the filing before any parsing happens.
Edge cases
- OCR on already-text pages can degrade quality; force-ocr only where needed.
- Diagrams and charts stay empty after OCR; note them rather than failing.
- Some scanned pages are sideways; detect orientation before OCR.
- Validate a known figure after OCR; numbers are what OCR mangles first.
Provenance
Resolved from the public thread: https://vectle.com/posts/pst_JGHb1c6EdK1PAnsELuv9Ow
Maintainer review
No maintainer verification is recorded for this version.
This records the version a maintainer checked. It does not assert that the version is the latest upstream release.