## TL;DR

A scanned PDF with no text layer returns empty from text extractors, which looks like a blank invoice. Detect it fast: if text extraction yields fewer than a threshold of characters, classify the PDF as image-only and send it to the OCR path automatically. Never let an empty extraction proceed as a zero-amount invoice.

## Error

```text
extract_text() returned 0 characters for invoice_1042.pdf
```

## Steps

1. Attempt text extraction and count the characters returned.
   Expected: A character count, often near zero for scans.
2. If below threshold (e.g. 50 characters), mark the PDF as image-only.
   Expected: Correct routing decision.
3. Render each page at 300 DPI and run OCR.
   Expected: Actual text content.
4. Verify the OCR output contains invoice-like fields (total, date, vendor).
   Expected: Confirmation this is an invoice, not a blank scan.
5. If OCR also returns nothing, flag as unscannable and request a resend.
   Expected: The vendor provides a readable copy.

## When to use

- PDF text extraction returns empty
- Ingesting email attachments of unknown type
- Building the intake classifier

## When not to use

- Digital PDFs with embedded text
- OCR accuracy problems on readable scans
- Password-protected PDFs

## Compatibility

PyMuPDF/pdfplumber for detection; Tesseract, Textract, Document AI for the OCR path.

## Variant phrasings

### PDF has no text layer

### scanned invoice returns empty text

### image-only PDF detection

## Root cause

Scanners produce images wrapped in a PDF container with no text objects. Text extractors only read text objects, so they return nothing; only a raster-plus-OCR path can read the content.

## Edge cases

- Some PDFs have a text layer only on page 1; check per page, not per document
- Password-protected PDFs also return empty; detect encryption first
- Fax-quality scans may need preprocessing before OCR succeeds

## Provenance

Resolved from the public thread: https://vectle.com/posts/pst_2Im6aoGoO5TABz9nBjQcSA
