pdfminer returns empty text on scanned annual report, parsing failed
This skill fixes pdfminer returning empty text on scanned annual reports. Use it when extraction yields nothing or when deciding if a PDF needs OCR. It is not for garbled real text, which is an encoding issue; the fix is detecting scanned pages, running OCR, and validating the result before parsing.
pdfminer returns empty text on a scanned annual report
TL;DR
pdfminer returns nothing on scanned annual reports because there is no text to extract, only images of pages, and no text extractor can fix that. Detect the scanned pages first, then run OCR to create a text layer before any parsing. After OCR, treat the document like any other PDF, but expect lower accuracy on tables and footnotes and validate more aggressively.
The error
(empty output, no exception)
pdfminer extracted 0 characters from annual report PDF; parsing failedWhen this helps
- pdfminer or any extractor returns empty text on a PDF
- annual reports or older filings yield nothing to parse
- deciding whether a PDF needs OCR
- validating OCR quality before table extraction
When it doesn't
- the PDF has real text but extraction is garbled; that is an encoding problem, not scanning
- you need perfect tables; OCR tables need heavy validation
- the scan is illegible; no OCR fixes a bad scan, find a better copy
Works with
python 3.8+ with pdfminer.six and pymupdf; ocrmypdf with tesseract. OCR quality depends on scan DPI.
Steps
1. Confirm the pages are scanned images, not text
from pdfminer.high_level import extract_text
text = extract_text("annual.pdf", maxpages=2)
print("chars extracted:", len(text.strip()))
import fitz
doc = fitz.open("annual.pdf")
print("images on page 1:", len(doc[0].get_images()))Expected: Near-zero characters plus images on the page. That combination is definitive: the PDF is scanned and needs OCR.
2. Run OCR to build a text layer
ocrmypdf --force-ocr annual.pdf annual_ocr.pdf
ls -la annual_ocr.pdfExpected: An OCRed PDF with a real text layer. Force-ocr handles pages that mix scanned images with a broken text layer.
3. Re-extract text from the OCRed file
from pdfminer.high_level import extract_text
text = extract_text("annual_ocr.pdf", maxpages=2)
print("chars after OCR:", len(text.strip()))
print(text[:300])Expected: Thousands of characters. If the count is still near zero, the scan quality is too poor and the source needs a better copy.
4. Validate OCR quality on a known figure
from pdfminer.high_level import extract_text
text = extract_text("annual_ocr.pdf")
import re
print("revenue mentions:", len(re.findall(r"revenue", text, re.I)))
print("year mentions:", len(re.findall(r"2026", text)))Expected: Sane hit counts for expected terms. OCR mangles numbers first, so spot-check a known figure from the report before trusting tables.
Other ways people phrase this
pdfminer empty text scanned pdf
The definitive symptom. Zero characters plus page images equals scanned.
pdf text extraction returns nothing annual report
Older annual reports are often scanned. OCR is the standard fix.
ocrmypdf annual report parsing
OCR-then-extract is the pipeline. Validate the OCR before trusting downstream tables.
Why it happens
Scanned PDFs contain page images with no text objects, so text extractors correctly return nothing: there is no text in the file. OCR adds a text layer by recognizing glyphs in the images. The empty output is not a pdfminer bug; it is the accurate report that the file has no extractable text.
Edge cases
- 300 DPI scans OCR far better than 150 DPI; if you control the scan, rescan at higher resolution.
- OCR confuses 0 with O and 1 with l in financial tables; validate numbers against known totals.
- Some PDFs mix scanned and born-digital pages; OCR only the pages that need it to save time.
- Password-protected scans need unlocking before OCR; that is a separate step.
Provenance
Resolved from the public thread: https://vectle.com/posts/pst_zL0gdAK0BaSbS0vnxwgvhw
Maintainer review
No maintainer verification is recorded for this version.
This records the version a maintainer checked. It does not assert that the version is the latest upstream release.