## TL;DR

Duplicates with retyped numbers or slightly different amounts defeat exact matching. Score candidate pairs on amount similarity, date proximity, vendor similarity, and line-item overlap; route pairs above threshold to review with a side-by-side diff. Tune the threshold on labeled historical data URIs start high and lower it until the review queue fills with real catches, not noise.

## Steps

1. Generate candidate pairs within a date window and amount band.
   Expected: A bounded candidate set.
2. Score pairs on amount, date, vendor, and line overlap.
   Expected: Ranked suspicions.
3. Route above-threshold pairs to review with diffs.
   Expected: Human decisions on the best candidates.
4. Label outcomes to build training data.
   Expected: A labeled set that compounds.
5. Retune thresholds quarterly.
   Expected: Precision that holds over time.

## When to use

- Retyped invoice numbers
- Slight amount variations
- Mature AP control programs

## When not to use

- Exact duplicates (simpler check)
- Real-time posting (batch this)
- Single-vendor shops

## Compatibility

Python (rapidfuzz, sklearn); ERP-agnostic.

## Variant phrasings

### near duplicate invoice

### fuzzy duplicate detection AP

### similar invoice detection

## Root cause

Real duplicates are rarely byte-identical: numbers get retyped, amounts gain fees, dates shift. Fuzzy scoring catches the family resemblance that exact keys miss.

## Edge cases

- Batch the scoring; pairwise comparison is O(n^2) without windowing
- Recurring invoices need allowlisting or they dominate the flags
- Thresholds drift as volume grows; retune on schedule

## Provenance

Resolved from the public thread: https://vectle.com/posts/pst_3-m5TP7JkYHwdLnid_kcqg
