annual report pdf text extraction unicode garbled, extraction failed
This skill fixes garbled unicode from annual report PDF text extraction. Use it when extractors return mojibake or null-byte text, or when choosing an extractor. It is not for scanned PDFs; the fix is diagnosing the encoding damage, trying a more forgiving extractor, repairing null bytes, and falling back to OCR.
Annual report PDF text extraction comes out as garbled unicode
TL;DR
Garbled unicode from annual report PDFs means the file's font encoding maps glyphs to wrong code points, so the extractor faithfully returns the wrong characters. The layout is fine; the character map is lying. Fix it by extracting with a tool that handles the encoding better, falling back to OCR when the map is unrecoverable, and validating with a known string from the report before trusting the output.
The error
(garbled output, no exception)
annual report text extraction returned mojibake: "R[NUL]e[NUL]v[NUL]e[NUL]n[NUL]u[NUL]e" / wrong glyphs throughoutWhen this helps
- PDF text extraction returns mojibake or null-byte text
- annual report parsing yields wrong characters
- choosing a text extractor for a new pipeline
- validating extraction quality before analysis
When it doesn't
- the PDF is scanned images; that needs OCR from the start, not encoding repair
- text is fine but tables are broken; that is a table extraction problem
- only some symbols are wrong; that may be a font subsetting quirk, not garbling
Works with
python 3.8+ with pdfminer.six, pymupdf, or pdfplumber; ocrmypdf with tesseract as fallback.
Steps
1. Identify the encoding damage pattern
from pdfminer.high_level import extract_text
text = extract_text("annual.pdf", maxpages=1)
print(repr(text[:200]))Expected: The raw repr. Null bytes between characters mean UTF-16 read as Latin-1; consistent wrong glyphs mean a broken ToUnicode map.
2. Try a different extractor with better encoding handling
import fitz
doc = fitz.open("annual.pdf")
text = doc[0].get_text()
print(repr(text[:200]))
print("chars:", len(text))Expected: Clean text or a different garble pattern. PyMuPDF and pdfplumber handle font encodings differently; one often succeeds where another fails.
3. Repair the common null-byte pattern when present
from pdfminer.high_level import extract_text
text = extract_text("annual.pdf", maxpages=1)
NUL = chr(0)
if NUL in text:
text = text.replace(NUL, "")
print(repr(text[:200]))Expected: Readable text. Stripping null bytes fixes the UTF-16-as-Latin-1 misread, which is the most common garble in annual reports.
4. Fall back to OCR when the font map is unrecoverable
ocrmypdf --force-ocr annual.pdf annual_ocr.pdf
python3 -c "from pdfminer.high_level import extract_text; t=extract_text('annual_ocr.pdf', maxpages=1); print(repr(t[:200]))"Expected: Clean OCR text. When the ToUnicode map is missing entirely, no extractor can recover the text and OCR is the only fix.
Other ways people phrase this
pdf text extraction unicode garbled
The general symptom. Check the repr first; the damage pattern picks the fix.
annual report pdf mojibake text
Font encoding issue, not a corrupt file. A different extractor or null-byte repair usually fixes it.
pdfminer wrong characters extraction
pdfminer is strict about font maps. PyMuPDF is more forgiving; try it second.
Why it happens
PDFs map glyphs to characters through font encoding tables, and many annual report producers emit broken or missing ToUnicode maps. Extractors read the bytes faithfully and produce the wrong characters because the map lies. The visual rendering is fine because it uses glyph shapes, not character codes, which is why the PDF looks right and extracts wrong.
Edge cases
- Ligatures like fi can extract as single private-use characters; normalize them after extraction.
- Some CJK or symbol fonts have no unicode map at all; OCR is the only recovery.
- Repairing null bytes is safe, but do not guess-replace other wrong glyphs without validation.
- Validate with a known phrase from the report, like the company name, before trusting numbers.
Provenance
Resolved from the public thread: https://vectle.com/posts/pstPs1Sd8SqlENAa5MIHoNNw
Maintainer review
No maintainer verification is recorded for this version.
This records the version a maintainer checked. It does not assert that the version is the latest upstream release.