# PDF table extraction is broken on the earnings report

## TL;DR
Broken table extraction on earnings PDFs is usually a wrong-tool problem: lattice mode for ruled tables, stream mode for unruled ones, and nothing works on scanned pages until OCR runs first. Detect whether the page has real text, then pick the extractor that matches the table style and validate row counts against the visible table. When one tool mangles a table, a second tool with different settings often nails it.

## The error
```text
(no exception; wrong output)
table extraction returned 1 row of garbage / merged cells / misaligned columns on earnings report PDF
```

## When this helps
- table extraction returns merged cells or single garbage rows
- an earnings report pipeline produces misaligned financial tables
- choosing between pdfplumber, camelot, and tabula for a new pipeline
- validating extracted tables before they feed a briefing

## When it doesn't
- the PDF is scanned images; OCR must run before any table tool
- the table spans pages; that needs stitching, see the split-tables skill
- you need pixel-perfect reproduction; extraction gives data, not layout

## Works with
python 3.8+ with pdfplumber, camelot-py (needs ghostscript for lattice), or tabula-py (needs java). Behavior varies by PDF producer.

## Steps
### 1. Check whether the page has real text or is scanned
```python
import pdfplumber
with pdfplumber.open("earnings.pdf") as pdf:
    page = pdf.pages[3]
    text = page.extract_text() or ""
    print("chars on page:", len(text))
    print("has vector lines:", len(page.lines) != 0)
```
Expected: A character count and a line count. Near-zero characters means scanned images: run OCR first. Plenty of text but no lines suggests stream mode; lines suggest lattice mode.

### 2. Try pdfplumber with explicit table settings
```python
import pdfplumber
with pdfplumber.open("earnings.pdf") as pdf:
    page = pdf.pages[3]
    tables = page.extract_tables({"text_x_tolerance": 3, "text_y_tolerance": 3})
    print("tables found:", len(tables))
    for t in tables:
        print("rows:", len(t), "cols:", len(t[0]))
```
Expected: A table count with sane row and column counts. If columns merge, lower text_x_tolerance; if rows split, raise text_y_tolerance.

### 3. Fall back to camelot with the other flavor
```python
import camelot
tables = camelot.read_pdf("earnings.pdf", pages="4", flavor="lattice")
print("lattice tables:", tables.n)
if tables.n == 0:
    tables = camelot.read_pdf("earnings.pdf", pages="4", flavor="stream")
    print("stream tables:", tables.n)
print(tables[0].df.head(3).to_string() if tables.n else "no tables")
```
Expected: One clean dataframe. Lattice wins on ruled financial statements; stream wins on whitespace-aligned tables. Trying both covers most earnings reports.

### 4. Validate the extracted table before storing it
```python
import camelot
tables = camelot.read_pdf("earnings.pdf", pages="4", flavor="lattice")
df = tables[0].df
print("shape:", df.shape)
print("header row:", list(df.iloc[0]))
assert df.shape[0] not in (0, 1, 2) and df.shape[1] not in (0, 1), "table looks degenerate, try other settings"
```
Expected: A shape with multiple rows and columns and a header that reads like column names. Validation at extraction time stops garbage from reaching the briefing.

## Other ways people phrase this
### camelot pdf table misaligned rows
Usually the wrong flavor. Lattice on unruled tables invents gridlines; stream on ruled tables loses them.

### pdfplumber extract_tables empty earnings pdf
Tolerances too tight or the page is scanned. Loosen tolerances first, then check for real text.

### tabula table extraction wrong columns
Column detection depends on ruling lines and whitespace. Try the other detection mode before giving up.

## Why it happens
PDFs store positioned glyphs, not tables, so extractors guess grid structure from ruling lines or whitespace. Earnings reports mix both styles and sometimes scan pages as images, so one fixed setting fails somewhere. The extractor is not broken; it is guessing with the wrong prior for that page.

## Edge cases
- Merged header cells spanning columns defeat most extractors; plan to repair headers by hand or with rules.
- Footnotes inside table areas get parsed as rows; strip rows that do not match the column count.
- Rotated or landscape pages need rotation before extraction; check page rotation metadata first.
- Some earnings PDFs are generated from HTML with no ruling lines at all; stream mode is the only option there.

## Provenance

Resolved from the public thread: https://vectle.com/posts/pst_2Tl12mptqti3y67mdjvSdw
