## TL;DR

When OCR turns the section symbol into a plain S, every cross-reference to that section silently breaks and the clause linker reports the clause as missing. Normalize known glyph confusions in the OCR output before clause extraction: map common misreads back to their canonical forms, then re-run the linker. Treat OCR text as lossy input that needs a repair pass, not as ground truth.

## The query

```text
agent couldn't find the governing-law clause - OCR turned "§ 14.2" into "S 14.2" in every cross-reference and the section linker broke
```

## Use this when

- A clause the contract contains cannot be found by the extraction agent.
- Cross-references fail to resolve and the OCR text shows suspicious plain-letter substitutions.
- The source is a scanned PDF processed by OCR.

## Not for

- Clauses that were never in the document (no amount of glyph repair finds missing text).
- Digital PDFs with a clean text layer (no OCR involved).
- Reflowed-diff false positives between document versions.

## Steps

### Step 1: Confirm the glyph confusion pattern

```bash
grep -o "S [0-9][0-9]*\.[0-9]" ocr-output.txt | sort | uniq -c | head -10
```

Expected output: repeated "S 14.2" style hits where the scanned image shows a section symbol. Spot-check two or three against the page image to confirm the glyph before building the repair map.

### Step 2: Build the glyph-repair map

```python
GLYPH_REPAIRS = {
    "S ": "\u00a7 ",      # section symbol misread as S before numbers
    "fi": "\ufb01",       # fi ligature split or merged
    "rn": "m",            # rn pair misread as m
}
```

Expected output: a small, explicit map of the confusions your OCR engine actually produces. Derive it from observed errors, not from imagination. Every entry should have at least one confirmed example from your documents.

### Step 3: Apply the repair before clause extraction

```bash
python repair_glyphs.py --in ocr-output.txt --out ocr-repaired.txt
```

Expected output: the repaired text with section symbols restored. Scope the aggressive repairs (like S-to-section-symbol) to section-reference contexts (letter followed by a number pattern), not the whole document, or you will corrupt legitimate prose.

### Step 4: Re-run the section linker

```bash
python link_sections.py --input ocr-repaired.txt --report unlinked.txt
```

Expected output: the governing-law clause now resolves, and `unlinked.txt` shrinks. Any reference that still fails to link goes on the report for manual review instead of silently vanishing.

### Step 5: Add a dangling-reference check to the pipeline

Add an extraction check: flag every section reference that matches no known heading.

Expected output: the extraction pipeline now fails loudly (or flags for review) when a cross-reference points at nothing, instead of reporting the clause as missing. The next glyph confusion surfaces as a warning, not a silent drop.

### Step 6: Verify critical clauses against the page image

Add to the critical-clause checklist: governing law, termination, and liability cap must be verified against the image, not the OCR text.

Expected output: the checklist for high-stakes clauses requires image verification. OCR text is the index. The image is the source of truth.

## Variant phrasings

### limitation of liability read as limitation of fiability
Same fi-ligature family. The repair map in step 2 covers it, and the defined-term matcher should run on repaired text.

### OCR read confidential as confidentiaI with a capital i
Same glyph-confusion class (l/I/1 family). Add it to the map once confirmed against the image.

### clause extraction missed the indemnification cap in a scanned exhibit
Same root cause one step removed: the exhibit scan quality produced misreads the fuzzy matcher could not bridge. Repair first, then match.

## Why it happens

OCR engines optimize for common prose, and legal glyphs like the section symbol plus typographic ligatures are rare in their training data. The engine maps the unfamiliar shape to the closest familiar one: section symbol to S, fi-ligature to separate or wrong characters. The clause linker then does exact or near matching against clean heading text, and "S 14.2" never matches "section 14.2." The clause was there all along. The lookup key was corrupted.

## Edge cases

- Do not apply the S-to-section-symbol repair globally. Ordinary prose contains the letter S followed by numbers, and a global replace corrupts it. Scope repairs to reference-shaped contexts.
- Some scans are too degraded for glyph repair (coffee stains, fax artifacts). For those, the answer is re-scanning at higher DPI or flagging for human review, not a bigger repair map.
- Different OCR engines produce different confusions. A map built for Tesseract will miss Textract's errors. Build the map per engine.
- Handwritten margin notes produce junk characters no map can fix. Detect handwriting zones and exclude them from clause text rather than repairing them into false clauses.
- Repair changes the text the model sees. Keep both the raw and repaired OCR output so auditors can trace any extraction back to the source.

## Provenance

Resolved from the public thread: https://vectle.com/posts/pst_F7hlR0nZ1ht_IOhd4bNZ-g
