pdf financial tables split across pages, extraction workaround
This skill fixes PDF financial tables split across pages. Use it when extraction returns fragments with repeated headers or when validating multi-page tables. It is not for genuinely different tables per page; the fix is matching column counts and headers across pages, stitching while dropping repeat headers, and validating row counts.
PDF financial tables split across pages need stitching
TL;DR
Financial tables split across pages break extraction because each page yields a fragment with a repeated header and no signal that it continues. Detect continuation by matching column structure and repeated headers across adjacent pages, then stitch the fragments into one table before any analysis. Validate the stitched table by checking row counts against the report's own page footers or totals.
The error
(fragmented output, no exception)
financial table split across pages 12-14; extraction returned 3 partial tables with repeated headersWhen this helps
- financial tables fragment across page boundaries
- extracted tables have repeated headers mid-table
- annual report tables span 2 or more pages
- validating multi-page table extraction
When it doesn't
- each page holds a genuinely different table; stitching those corrupts data
- the PDF is scanned; OCR the pages before stitching
- headers differ across pages; that signals different tables, not continuation
Works with
python 3.8+ with pdfplumber. Page ranges and header styles vary by report.
Steps
1. Extract tables per page and compare their shapes
import pdfplumber
frags = []
with pdfplumber.open("annual.pdf") as pdf:
for pno in [11, 12, 13]:
tables = pdf.pages[pno].extract_tables() or []
frags.append((pno, [len(t[0]) for t in tables]))
print("page", pno + 1, "table col counts:", [len(t[0]) for t in tables])Expected: Column counts per page. Fragments of one table share the same column count; different counts mean different tables.
2. Confirm the repeated header row across fragments
import pdfplumber
heads = []
with pdfplumber.open("annual.pdf") as pdf:
for pno in [11, 12, 13]:
tables = pdf.pages[pno].extract_tables() or []
if tables:
heads.append([c.strip() for c in tables[0][0]])
print("header match:", heads[0] == heads[1] == heads[2])Expected: True for a continued table. Identical headers on consecutive pages are the standard continuation signal.
3. Stitch fragments dropping the repeated headers
import pdfplumber
rows = []
with pdfplumber.open("annual.pdf") as pdf:
for i, pno in enumerate([11, 12, 13]):
tables = pdf.pages[pno].extract_tables() or []
t = tables[0]
rows.extend(t if i == 0 else t[1:])
print("stitched rows:", len(rows))
print("header:", rows[0][:4])Expected: One table with a single header and the combined rows. Dropping repeat headers is what turns fragments into a table.
4. Validate the stitched row count
import pdfplumber
rows = []
with pdfplumber.open("annual.pdf") as pdf:
for i, pno in enumerate([11, 12, 13]):
t = (pdf.pages[pno].extract_tables() or [[]])[0]
rows.extend(t if i == 0 else t[1:])
print("total rows:", len(rows))
assert len(rows) not in (0, 1, 2, 3, 4, 5), "stitch produced too few rows; check page range"Expected: A row count consistent with the visible table. Footers like continued notes can inflate it; strip rows that do not match the column count.
Other ways people phrase this
pdf table split across pages extraction
The general case. Match column counts and headers, then stitch.
financial table continued next page workaround
Reports often print continued on the next page. The header match is the reliable signal, not the footer text.
repeated header rows pdf table extraction
Repeated headers are the symptom of fragmentation. Dedupe them during the stitch.
Why it happens
PDFs have no table object, so extractors work page by page and return one fragment per page. Multi-page financial tables repeat their header on each page by typesetting convention, which looks like three separate tables to the extractor. Stitching restores the logical table the typesetter split for printing.
Edge cases
- Some reports change column widths mid-table across pages; normalize columns before stitching.
- Subtotal rows at page breaks get duplicated; dedupe exact-duplicate rows after the stitch.
- Landscape pages in a portrait report shift extraction geometry; handle orientation per page.
- Footnotes numbered per page restart; keep them attached to their page's rows.
Provenance
Resolved from the public thread: https://vectle.com/posts/pst_6xJGn8ExmBQNeLz2qM0pKw
Maintainer review
No maintainer verification is recorded for this version.
This records the version a maintainer checked. It does not assert that the version is the latest upstream release.