how to handle line items split across pages in invoice PDFs
Merges invoice line-item rows that continue across page boundaries into single logical rows. Use when a multi-page invoice table splits rows mid-row or repeats headers on each page. Not for single-page invoices or header-field extraction.
TL;DR
Multi-page invoices split table rows across pages and repeat header rows, which naive extractors read as duplicate or broken line items. Extract each page separately with explicit page numbers, drop repeated header rows, then stitch rows by matching column counts and checking that quantities and amounts parse. Validate by confirming row count and the sum of extended prices against the subtotal.
Steps
- Extract each page independently, tagging every row with its page number.
Expected: Per-page row lists with page annotations.
- Detect and drop repeated header rows by matching the header text pattern on pages 2+.
Expected: Header rows removed, only data rows remain.
- Stitch split rows: if the last row of page N has fewer populated columns than the header, merge it with the first row of page N+1.
Expected: Complete logical rows spanning the page break.
- Recompute subtotal from merged rows and compare to the invoice subtotal.
Expected: Match within rounding tolerance confirms the stitch is correct.
- Flag any row that still fails column-count validation for human review.
Expected: Only genuinely ambiguous rows need a person.
When to use
- Invoices longer than one page with itemized tables
- Extracted line items show partial rows or duplicated headers
- Subtotal does not reconcile with extracted rows
When not to use
- Single-page invoices
- Header fields like vendor or date (not tables)
- Digital invoices with structured line-item XML
Compatibility
pdfplumber, PyMuPDF, AWS Textract (with page numbers), Google Document AI. Works on scanned and digital PDFs.
Variant phrasings
invoice table continues on next page
merge line items across PDF pages
multi-page invoice row stitching
Root cause
PDF has no table-row concept across pages; each page is an independent layout. Extractors process pages in isolation, so a row that starts at the bottom of page 1 and finishes at the top of page 2 becomes two fragments, and repeated column headers look like data.
Edge cases
- Footers with page totals mid-table can be mistaken for line items; filter rows matching total patterns
- Landscape pages mixed into a portrait invoice shift column x-positions; normalize per page
- Some vendors restart row numbering per page; do not use row numbers as merge keys
Provenance
Resolved from the public thread: https://vectle.com/posts/pst_080tcvVn4gz3zBIbnUU-tw
Maintainer review
No maintainer verification is recorded for this version.
This records the version a maintainer checked. It does not assert that the version is the latest upstream release.