TL;DR
OCR engines confuse the lowercase l with a capital I, so "confidential" comes out as "confidentiaI" and an exact-match detector reports no NDA present. Normalize the known confusion pairs (l, I, 1 and friends) in a post-OCR pass, then match keywords with fuzzy scoring instead of exact equality. The detector finds the term even when individual glyphs are wrong.

The misread text in your OCR output:
```text
confidentiaI
```

## Steps
1. Confirm the term is present but misspelled in your OCR text. Search for near-misses of the keyword.
   Command: grep -rni "confident" ocr_text/ | head
   Expected: hits showing "confidentiaI" or similar, proving the NDA language is there and the detector's exact match is the failure point.
2. Add a confusion-pair normalization pass right after OCR. Map the classic lookalike glyphs to one canonical form before any detection runs.
   ```python
   CONFUSIONS = {"I": "l", "1": "l", "0": "O"}
   def normalize_confusables(text):
       for wrong, right in CONFUSIONS.items():
           text = text.replace(wrong, right)
       return text
   ```
   Expected: "confidentiaI" becomes "confidential" in the normalized stream.
3. Replace exact keyword matching with fuzzy matching for presence checks. Score candidate words against your keyword list with a tolerance for one or two character substitutions.
   Expected: the detector reports the NDA clause present with a confidence above your presence threshold.
4. Re-run detection on the document and also grep for any remaining near-misses your map did not cover.
   Expected: no remaining near-miss spellings of "confidential" in the text, and the NDA flag is set.

## Use this when
- Scanned contracts where lowercase l, capital I, and digit 1 are visually identical in the font
- Exact-match detectors report a term absent when the page clearly contains it
- OCR output has words with a stray capital letter in the middle, like "confidentiaI" or "liaBiIity"
- Keyword presence checks gate downstream logic like NDA detection or DPA detection

## Not for this skill when
- The document is born-digital, read the text layer instead of fighting OCR
- The term is genuinely absent, check the unnormalized text before assuming a glyph error
- Your detector already uses embeddings or fuzzy matching, look at your match threshold instead
- The confusion is between unrelated shapes, like rn read as m, which needs a different confusion map

## Variant phrasings
- OCR read confidential as confidentiaI, detector said no NDA
- tesseract l versus I confusion in scanned contracts
- exact match missed the keyword because of OCR glyph errors
- scanned NDA not detected, lowercase L read as capital i
- OCR confusables breaking clause presence checks

## Why it happens
In many legal fonts the lowercase l, capital I, and digit 1 are nearly identical vertical strokes, and at scan resolution there is not enough pixel detail to tell them apart. The OCR engine picks one, often the capital I, and your downstream code does a byte-exact comparison against "confidential", which fails on that single character. Exact matching is the wrong tool for OCR text because OCR output is approximate by nature, one wrong glyph per hundred characters is a normal error rate.

## Edge cases
- The "0" to "O" map can corrupt real numbers, apply the full confusion map only to alphabetic keyword matching, and keep a separate digit-safe pass for amounts and dates.
- Some fonts genuinely use a capital I inside a word, like brand names, so run fuzzy matching on the normalized text but quote the original text in citations.
- Fuzzy matching thresholds need tuning per keyword length, short words like "term" need a tighter threshold than "confidentiality".
- If your OCR engine reports per-character confidences, use them to prefer the low-confidence character's alternatives instead of a blind global replace.
- Watch for the reverse error too, a real digit 1 inside a clause number becoming lowercase l, your map handles both directions.

## Provenance

Resolved from the public thread: https://vectle.com/posts/pst_dY5TOrQHHlyeHeyZNpZ15g
