pdf table extraction broken on earnings report, how to fix
This skill fixes broken PDF table extraction on earnings reports. Use it when extractors return merged cells, garbage rows, or empty tables, or when picking a table tool. It is not for scanned PDFs, which need OCR first; the fix is detecting the table style, matching the extractor flavor to it, and validating row counts.
PDF table extraction is broken on the earnings report
TL;DR
Broken table extraction on earnings PDFs is usually a wrong-tool problem: lattice mode for ruled tables, stream mode for unruled ones, and nothing works on scanned pages until OCR runs first. Detect whether the page has real text, then pick the extractor that matches the table style and validate row counts against the visible table. When one tool mangles a table, a second tool with different settings often nails it.
The error
(no exception; wrong output)
table extraction returned 1 row of garbage / merged cells / misaligned columns on earnings report PDFWhen this helps
- table extraction returns merged cells or single garbage rows
- an earnings report pipeline produces misaligned financial tables
- choosing between pdfplumber, camelot, and tabula for a new pipeline
- validating extracted tables before they feed a briefing
When it doesn't
- the PDF is scanned images; OCR must run before any table tool
- the table spans pages; that needs stitching, see the split-tables skill
- you need pixel-perfect reproduction; extraction gives data, not layout
Works with
python 3.8+ with pdfplumber, camelot-py (needs ghostscript for lattice), or tabula-py (needs java). Behavior varies by PDF producer.
Steps
1. Check whether the page has real text or is scanned
import pdfplumber
with pdfplumber.open("earnings.pdf") as pdf:
page = pdf.pages[3]
text = page.extract_text() or ""
print("chars on page:", len(text))
print("has vector lines:", len(page.lines) != 0)Expected: A character count and a line count. Near-zero characters means scanned images: run OCR first. Plenty of text but no lines suggests stream mode; lines suggest lattice mode.
2. Try pdfplumber with explicit table settings
import pdfplumber
with pdfplumber.open("earnings.pdf") as pdf:
page = pdf.pages[3]
tables = page.extract_tables({"text_x_tolerance": 3, "text_y_tolerance": 3})
print("tables found:", len(tables))
for t in tables:
print("rows:", len(t), "cols:", len(t[0]))Expected: A table count with sane row and column counts. If columns merge, lower textxtolerance; if rows split, raise textytolerance.
3. Fall back to camelot with the other flavor
import camelot
tables = camelot.read_pdf("earnings.pdf", pages="4", flavor="lattice")
print("lattice tables:", tables.n)
if tables.n == 0:
tables = camelot.read_pdf("earnings.pdf", pages="4", flavor="stream")
print("stream tables:", tables.n)
print(tables[0].df.head(3).to_string() if tables.n else "no tables")Expected: One clean dataframe. Lattice wins on ruled financial statements; stream wins on whitespace-aligned tables. Trying both covers most earnings reports.
4. Validate the extracted table before storing it
import camelot
tables = camelot.read_pdf("earnings.pdf", pages="4", flavor="lattice")
df = tables[0].df
print("shape:", df.shape)
print("header row:", list(df.iloc[0]))
assert df.shape[0] not in (0, 1, 2) and df.shape[1] not in (0, 1), "table looks degenerate, try other settings"Expected: A shape with multiple rows and columns and a header that reads like column names. Validation at extraction time stops garbage from reaching the briefing.
Other ways people phrase this
camelot pdf table misaligned rows
Usually the wrong flavor. Lattice on unruled tables invents gridlines; stream on ruled tables loses them.
pdfplumber extract_tables empty earnings pdf
Tolerances too tight or the page is scanned. Loosen tolerances first, then check for real text.
tabula table extraction wrong columns
Column detection depends on ruling lines and whitespace. Try the other detection mode before giving up.
Why it happens
PDFs store positioned glyphs, not tables, so extractors guess grid structure from ruling lines or whitespace. Earnings reports mix both styles and sometimes scan pages as images, so one fixed setting fails somewhere. The extractor is not broken; it is guessing with the wrong prior for that page.
Edge cases
- Merged header cells spanning columns defeat most extractors; plan to repair headers by hand or with rules.
- Footnotes inside table areas get parsed as rows; strip rows that do not match the column count.
- Rotated or landscape pages need rotation before extraction; check page rotation metadata first.
- Some earnings PDFs are generated from HTML with no ruling lines at all; stream mode is the only option there.
Provenance
Resolved from the public thread: https://vectle.com/posts/pst_2Tl12mptqti3y67mdjvSdw
Maintainer review
No maintainer verification is recorded for this version.
This records the version a maintainer checked. It does not assert that the version is the latest upstream release.