LLM extraction hallucinating line items: detection and prevention
Detects and prevents LLMs from fabricating plausible invoice line items when rows are occluded or split. Use when using a vision LLM for invoice extraction and line items cannot be verified. Not for pure OCR pipelines without an LLM step.
TL;DR
The dominant LLM failure on invoices is silent hallucination: plausible quantities and prices for rows that are occluded, split, or simply absent, emitted with full confidence. Prevent it with grounding: require every extracted line item to cite its source page and bounding region, then run arithmetic validation (qty times price equals extended, rows sum to subtotal). Reject any extraction where items lack citations or the math fails, and fall back to deterministic OCR plus rules.
Steps
- Prompt the model to return source page and bounding box for every line item.
Expected: Each item carries a citation to a region of the PDF.
- Validate arithmetic: quantity times unit price must equal extended price per row.
Expected: Hallucinated rows usually fail the math.
- Validate the roll-up: extended prices must sum to the subtotal.
Expected: A second independent check on the row set.
- Spot-check citations by cropping the cited region and re-reading it.
Expected: Ground truth for a sample of rows.
- On any failure, fall back to deterministic OCR plus rule-based parsing for that invoice.
Expected: No silent bad data reaches the ERP.
When to use
- Using a vision LLM to extract invoice line items
- Line-item accuracy matters for PO matching
- Invoices with occlusions, stamps, or split rows
When not to use
- Pure OCR pipelines with no LLM
- Header-only extraction
- Structured e-invoices (XML) where there is nothing to hallucinate
Compatibility
Any vision LLM (GPT-4o, Claude, Gemini) plus a PDF renderer for cropping. Pairs with Textract/Document AI as the fallback.
Variant phrasings
LLM invents invoice line items
vision model fabricates rows
grounding invoice extraction
Root cause
LLMs are trained to produce plausible completions, not to say 'I cannot see this row.' When a row is partially occluded or split across pages, the model fills the gap with a statistically likely row and reports it with the same confidence as a clearly seen one.
Edge cases
- Bounding boxes can themselves be hallucinated; the arithmetic check is the real gate
- Handwritten line additions are the hardest case; route them to human review by policy
- Token limits on long invoices cause truncation that looks like missing rows; extract page by page
Provenance
Resolved from the public thread: https://vectle.com/posts/pstYgaraeFsmSas0IpZ9LsHw
Maintainer review
No maintainer verification is recorded for this version.
This records the version a maintainer checked. It does not assert that the version is the latest upstream release.