# PDF financial tables split across pages need stitching

## TL;DR
Financial tables split across pages break extraction because each page yields a fragment with a repeated header and no signal that it continues. Detect continuation by matching column structure and repeated headers across adjacent pages, then stitch the fragments into one table before any analysis. Validate the stitched table by checking row counts against the report's own page footers or totals.

## The error
```text
(fragmented output, no exception)
financial table split across pages 12-14; extraction returned 3 partial tables with repeated headers
```

## When this helps
- financial tables fragment across page boundaries
- extracted tables have repeated headers mid-table
- annual report tables span 2 or more pages
- validating multi-page table extraction

## When it doesn't
- each page holds a genuinely different table; stitching those corrupts data
- the PDF is scanned; OCR the pages before stitching
- headers differ across pages; that signals different tables, not continuation

## Works with
python 3.8+ with pdfplumber. Page ranges and header styles vary by report.

## Steps
### 1. Extract tables per page and compare their shapes
```python
import pdfplumber
frags = []
with pdfplumber.open("annual.pdf") as pdf:
    for pno in [11, 12, 13]:
        tables = pdf.pages[pno].extract_tables() or []
        frags.append((pno, [len(t[0]) for t in tables]))
        print("page", pno + 1, "table col counts:", [len(t[0]) for t in tables])
```
Expected: Column counts per page. Fragments of one table share the same column count; different counts mean different tables.

### 2. Confirm the repeated header row across fragments
```python
import pdfplumber
heads = []
with pdfplumber.open("annual.pdf") as pdf:
    for pno in [11, 12, 13]:
        tables = pdf.pages[pno].extract_tables() or []
        if tables:
            heads.append([c.strip() for c in tables[0][0]])
print("header match:", heads[0] == heads[1] == heads[2])
```
Expected: True for a continued table. Identical headers on consecutive pages are the standard continuation signal.

### 3. Stitch fragments dropping the repeated headers
```python
import pdfplumber
rows = []
with pdfplumber.open("annual.pdf") as pdf:
    for i, pno in enumerate([11, 12, 13]):
        tables = pdf.pages[pno].extract_tables() or []
        t = tables[0]
        rows.extend(t if i == 0 else t[1:])
print("stitched rows:", len(rows))
print("header:", rows[0][:4])
```
Expected: One table with a single header and the combined rows. Dropping repeat headers is what turns fragments into a table.

### 4. Validate the stitched row count
```python
import pdfplumber
rows = []
with pdfplumber.open("annual.pdf") as pdf:
    for i, pno in enumerate([11, 12, 13]):
        t = (pdf.pages[pno].extract_tables() or [[]])[0]
        rows.extend(t if i == 0 else t[1:])
print("total rows:", len(rows))
assert len(rows) not in (0, 1, 2, 3, 4, 5), "stitch produced too few rows; check page range"
```
Expected: A row count consistent with the visible table. Footers like continued notes can inflate it; strip rows that do not match the column count.

## Other ways people phrase this
### pdf table split across pages extraction
The general case. Match column counts and headers, then stitch.

### financial table continued next page workaround
Reports often print continued on the next page. The header match is the reliable signal, not the footer text.

### repeated header rows pdf table extraction
Repeated headers are the symptom of fragmentation. Dedupe them during the stitch.

## Why it happens
PDFs have no table object, so extractors work page by page and return one fragment per page. Multi-page financial tables repeat their header on each page by typesetting convention, which looks like three separate tables to the extractor. Stitching restores the logical table the typesetter split for printing.

## Edge cases
- Some reports change column widths mid-table across pages; normalize columns before stitching.
- Subtotal rows at page breaks get duplicated; dedupe exact-duplicate rows after the stitch.
- Landscape pages in a portrait report shift extraction geometry; handle orientation per page.
- Footnotes numbered per page restart; keep them attached to their page's rows.

## Provenance

Resolved from the public thread: https://vectle.com/posts/pst_6xJGn8ExmBQNeLz2qM0pKw
