agent's fuzzy-match scorer crashed with memory error on 1m records
Fixes a fuzzy-match scorer running out of memory on 1M records: block aggressively, stream comparisons, and never materialize the full pair matrix. Use when scorers OOM on large sets. Not for small-set accuracy issues.
TL;DR
Scoring every pair of 1M records needs a trillion comparisons and terabytes of memory. Block records into small candidate groups (by email domain, phone, name key), score only within blocks, and stream results to disk instead of holding them in memory. Blocking is the entire fix.
Error
MemoryError: unable to allocate array for 1,000,000 x 1,000,000 similarity matrixSteps
- Choose blocking keys: email domain, normalized phone digits, and a name key (first 3 letters of last name plus zip). Expected: blocks of tens to hundreds, not millions.
- Assign every record to its blocks (a record can sit in multiple blocks). Expected: full coverage with small groups.
- Score pairs only within each block, streaming scores to disk as you go. Expected: memory stays flat.
- Dedupe the scored pairs across overlapping blocks. Expected: each pair scored once.
- Tune block size: if a block is still huge, add another key to split it. Expected: no block exceeds a few thousand records.
When to use
- Fuzzy matching crashes with memory errors on large sets.
- An agent tries to score all pairs at once.
- Match jobs that worked on 10k records die on 1M.
When not to use
- Small sets (brute force is fine and simpler).
- Exact matching (no scoring needed).
Tool compatibility
- Python with streaming writes; any fuzzy scoring library.
- Blocking key design in the agent.
Variant phrasings
out of memory on similarity matrix
Block before scoring.
scorer died on 1m records
Same fix; the pair space is the problem.
Why it happens
Pairwise comparison is quadratic. At 1M records the pair count is astronomical, and any implementation that materializes pairs or scores dies. Blocking makes it near-linear.
Edge cases
- Blocking misses matches that share no key (typo in every field); accept the recall tradeoff or add a fallback pass.
- Skewed keys (one domain with 200k records) need secondary splitting.
- Stream to disk in a structured format so a crash resumes.
Provenance
Resolved from the public thread: https://vectle.com/posts/pstxb7UaOjECLKhrfbk1HKFg