# Tabula and camelot misalign rows on PDF table extraction

## TL;DR
Misaligned rows from tabula and camelot mean the table's visual grid does not match the extractor's assumptions: multi-line cells, missing ruling lines, or merged headers shift everything. The fix is picking the right detection mode for the table's actual structure and post-processing the dataframe, not endlessly tuning one mode. When both tools misalign the same table, the table is genuinely irregular and needs rule-based cleanup.

## The error
```text
(misaligned output, no exception)
tabula/camelot rows shifted: values land in wrong columns, header merged into data rows
```

## When this helps
- tabula or camelot return misaligned rows
- values land in wrong columns after extraction
- troubleshooting table extraction on a new report
- deciding between lattice and stream modes

## When it doesn't
- the PDF is scanned; OCR first, then extract
- tables span pages; stitch fragments before aligning
- the table is an image inside the PDF; extraction cannot see it

## Works with
python 3.8+ with camelot-py (ghostscript for lattice) or tabula-py (java). Behavior varies by PDF producer.

## Steps
### 1. Compare lattice versus stream on the same page
```python
import camelot
for flavor in ["lattice", "stream"]:
    t = camelot.read_pdf("report.pdf", pages="5", flavor=flavor)
    print(flavor, t.n, "tables")
    if t.n:
        print(t[0].df.shape)
```
Expected: Shapes from both flavors. The flavor whose shape matches the visible table is the right starting point.

### 2. Inspect where the misalignment starts
```python
import camelot
t = camelot.read_pdf("report.pdf", pages="5", flavor="stream")
df = t[0].df
for i, row in df.head(8).iterrows():
    print(i, list(row)[:4])
```
Expected: The first shifted row. Misalignment usually starts at a multi-line cell or a merged header; everything after shifts.

### 3. Repair with column-aware cleanup rules
```python
import camelot, re
df = camelot.read_pdf("report.pdf", pages="5", flavor="stream")[0].df
df = df[df.apply(lambda r: sum(1 for c in r if str(c).strip()) != 0, axis=1)]
print("non-empty rows:", len(df))
print("cols:", df.shape[1])
```
Expected: A dataframe with empty artifact rows removed. Row filtering by non-empty cell count fixes the most common shift source.

### 4. Validate against a known total from the table
```python
import camelot
df = camelot.read_pdf("report.pdf", pages="5", flavor="stream")[0].df
txt = df.to_string()
print("mentions total:", "total" in txt.lower())
print("row count:", len(df))
```
Expected: Confirmation the totals row survived. If the total row is missing or shifted, the alignment is still wrong.

## Other ways people phrase this
### camelot rows misaligned pdf table
Usually multi-line cells or merged headers. Stream plus row cleanup handles most cases.

### tabula columns shifted extraction
Tabula's lattice mode invents columns from ruling lines; stream mode follows whitespace. Try both.

### pdf table values wrong columns
The alignment half of table extraction. Detect the shift start, then clean.

## Why it happens
Extractors assume tables are regular grids, but real financial tables have multi-line cells, merged headers, and inconsistent ruling. One irregular row shifts every row after it in stream mode, while lattice mode hallucinates gridlines from stray rules. The tools are not broken; the table violates their assumptions and needs cleanup rules.

## Edge cases
- Merged header cells never align cleanly; extract headers separately from body rows.
- Currency symbols in their own column defeat whitespace detection; strip them before aligning.
- Tables with both ruled and unruled sections need per-section settings.
- When both flavors fail identically, the table is irregular enough to warrant manual mapping.

## Provenance

Resolved from the public thread: https://vectle.com/posts/pst_fdhQuTh-K8MLmwnXgOELcg
