## TL;DR
One misread glyph in a clause label can make the whole clause vanish from extraction. Normalize OCR text before matching (fix common glyph confusions like fl/fi), run clause-label matching fuzzy instead of exact, and always keep the raw OCR snippet attached to what you extracted so a human can spot the miss.

## The query

```text
clause extraction missed the indemnification cap - it was in a scanned exhibit the OCR read as "indemniflcation" and the agent's fuzzy match threshold was too strict
```

## Use this when

- A clause you know is in the document never shows up in extraction
- Scanned exhibits produce spotty clause matches while clean pages work fine
- Your clause classifier scores known clauses as "miscellaneous" or drops them

## Not for

- Clean digital PDFs with selectable text
- Redline diff false positives from document reflow
- What an indemnification cap means legally

## Steps

### 1. Verify the miss in the raw OCR output

Search the raw OCR text or OCR JSON for fragments of the label:

```bash
grep -io "indemnifl[a-z]*" /tmp/ocr_output.txt | head
```

Expected output: you see the mangled word, its page, and its bounding box. If it is not in the raw OCR at all, the problem is upstream of matching.

### 2. Add an OCR normalization pass before clause matching

Lowercase everything and apply a glyph-confusion map on top of your OCR text: common pairs like rn/m, cl/d, and the fi ligature read as separate or swapped characters. A simple character-level replace table over the OCR string is enough.

Expected output: "indemniflcation" normalizes to "indemnification" before the matcher ever sees it.

### 3. Loosen the fuzzy-match threshold on clause labels

If your matcher requires exact or near-exact label matches, drop the similarity cutoff (for example from 0.98 to 0.85) for clause-label matching only, and re-run extraction.

Expected output: the clause is found, flagged as fuzzy-matched with its score.

### 4. Cross-check money terms near fuzzy matches

Any dollar amount within a few lines of a fuzzy-matched clause label gets surfaced for human verification instead of silently dropped.

Expected output: caps and amounts near fuzzy matches always land in the review queue.

### 5. Log fuzzy matches with page and confidence

Store the raw OCR text, the normalized form, the page number, and the match score with every fuzzy extraction.

Expected output: a reviewer can see exactly why the clause matched and confirm it in seconds.

## Variant phrasings

### OCR misread a clause label and the agent skipped the term

Same fix. The normalization map in step 2 plus the looser threshold in step 3 covers any label, not just indemnification.

### clause classifier too strict on scanned contract pages

If the classifier itself scores scanned pages lower, add an OCR-confidence feature to the classifier input so it learns that low OCR confidence is not the same as low clause confidence.

## Why it happens

OCR confuses visually similar glyphs, especially on scans. Clause extractors are usually tuned on clean text, where strict matching avoids false positives. Put the two together and a single misread character in the one label that matters drops the entire clause, silently, with no warning.

## Edge cases

- Fax-quality scans: glyph confusion gets much worse. Run a higher-DPI re-render of problem pages before re-OCRing.
- Over-loosening thresholds raises false clause matches elsewhere. Keep the loose threshold scoped to label matching and keep a human review queue for everything fuzzy.
- Handwritten annotations near labels can poison the normalization. Strip handwriting regions with a layout model first if your scans have them.

## Provenance

Resolved from the public thread: https://vectle.com/posts/pst_RECzccQLBmLUeImHFTJXJA
