scanned PDF with no text layer: extraction returns empty
Detects image-only scanned PDFs and routes them through OCR instead of text extraction. Use when a PDF yields no extractable text. Not for digital PDFs or OCR accuracy tuning.
TL;DR
A scanned PDF with no text layer returns empty from text extractors, which looks like a blank invoice. Detect it fast: if text extraction yields fewer than a threshold of characters, classify the PDF as image-only and send it to the OCR path automatically. Never let an empty extraction proceed as a zero-amount invoice.
Error
extract_text() returned 0 characters for invoice_1042.pdfSteps
- Attempt text extraction and count the characters returned.
Expected: A character count, often near zero for scans.
- If below threshold (e.g. 50 characters), mark the PDF as image-only.
Expected: Correct routing decision.
- Render each page at 300 DPI and run OCR.
Expected: Actual text content.
- Verify the OCR output contains invoice-like fields (total, date, vendor).
Expected: Confirmation this is an invoice, not a blank scan.
- If OCR also returns nothing, flag as unscannable and request a resend.
Expected: The vendor provides a readable copy.
When to use
- PDF text extraction returns empty
- Ingesting email attachments of unknown type
- Building the intake classifier
When not to use
- Digital PDFs with embedded text
- OCR accuracy problems on readable scans
- Password-protected PDFs
Compatibility
PyMuPDF/pdfplumber for detection; Tesseract, Textract, Document AI for the OCR path.
Variant phrasings
PDF has no text layer
scanned invoice returns empty text
image-only PDF detection
Root cause
Scanners produce images wrapped in a PDF container with no text objects. Text extractors only read text objects, so they return nothing; only a raster-plus-OCR path can read the content.
Edge cases
- Some PDFs have a text layer only on page 1; check per page, not per document
- Password-protected PDFs also return empty; detect encryption first
- Fax-quality scans may need preprocessing before OCR succeeds
Provenance
Resolved from the public thread: https://vectle.com/posts/pst_2Im6aoGoO5TABz9nBjQcSA
Maintainer review
No maintainer verification is recorded for this version.
This records the version a maintainer checked. It does not assert that the version is the latest upstream release.