duplicate line items extracted twice from one invoice
Dedupes line items that the extractor emitted twice from a single invoice. Use when row counts or subtotals look doubled. Not for duplicate invoices across documents.
TL;DR
Extractors double-emit rows when a table is processed twice (overlapping crops, repeated table detection) or when a wrapped description becomes its own row with copied amounts. Dedupe within the document: exact duplicate rows collapse, and rows whose amounts would double-count the subtotal get flagged. The subtotal reconciliation is the detector.
Error
Extracted 24 line items; invoice shows 12 (subtotal exactly 2x expected)Steps
- Hash each extracted row on description, quantity, and amounts.
Expected: Duplicate hashes reveal double-emitted rows.
- Collapse exact duplicates, keeping one copy.
Expected: A clean row list.
- Recompute the subtotal and compare to the invoice subtotal.
Expected: A match confirms the dedupe was correct.
- If the subtotal is still a multiple of expected, look for wrapped descriptions copied with amounts.
Expected: The subtler double-count pattern.
- Log the duplication pattern per extractor and vendor.
Expected: Feeds extractor tuning.
When to use
- Row count is a multiple of the visible row count
- Subtotal is 2x the invoice subtotal
- After table reprocessing or multi-pass extraction
When not to use
- Duplicate invoices (different documents)
- Genuinely repeated line items (same product ordered twice)
- PO-level duplicates
Compatibility
Extractor-agnostic; implement in the post-extraction normalization layer.
Variant phrasings
line items extracted twice
doubled rows invoice OCR
duplicate rows single invoice
Root cause
Table detectors can fire twice on the same table (overlapping page crops, repeated detection passes), and wrapped descriptions sometimes inherit the parent row's amounts. Both produce identical or near-identical rows.
Edge cases
- Genuinely duplicated lines (two identical items) must not be collapsed; use subtotal reconciliation as the arbiter
- Near-duplicates with OCR noise in descriptions need fuzzy hashing
- Multi-page tables need dedupe after page stitching, not before
Provenance
Resolved from the public thread: https://vectle.com/posts/pst_CtVghX4eD2C8ayzxTjqzYQ