TL;DR
Scanned fonts render "fi" as a single glyph, and OCR engines map that glyph to the wrong characters, so "limitation of liability" comes out as "limitation of fiability" and an exact-match classifier never fires. Normalize ligature codepoints immediately after OCR, then match clause headings with fuzzy scoring instead of exact string equality. The clause gets found even when the glyph is mangled.

The misread text in your OCR output:
```text
limitation of fiability
```

## Steps
1. Confirm the clause text is actually in your OCR output, just mangled. Run a search for the mangled heading across your extracted text.
   Command: grep -rni "fiability" ocr_text/
   Expected: a match on the indemnity page, proving the clause exists in the text and the classifier's exact match is what failed.
2. Add a ligature normalization pass immediately after OCR, before anything classifies the text. Some engines emit the real ligature codepoints, which then fail downstream string matches.
   ```python
   LIGATURES = {"\ufb01": "fi", "\ufb02": "fl", "\ufb00": "ff", "\ufb03": "ffi", "\ufb04": "ffl"}
   def normalize_ligatures(text):
       for glyph, plain in LIGATURES.items():
           text = text.replace(glyph, plain)
       return text
   ```
   Expected: any real ligature glyphs in the text become plain ASCII before classification runs.
3. Replace exact heading matching with fuzzy matching against your canonical heading list. Use a token-set scorer with a threshold around 88 to 90, so "limitation of fiability" still matches "limitation of liability".
   Expected: the classifier labels the clause as limitation-of-liability with a score above your threshold instead of dropping it to miscellaneous.
4. Re-run extraction on the document and check the previously missed clause.
   Expected: the indemnity clause appears in the output with its cap amount, and no other heading changes label.

## Use this when
- Scanned PDFs where fi, fl, or ff render as single ligature glyphs in the font
- A clause classifier misses a heading that looks almost right in the OCR text
- OCR output contains odd words like "fiability", "oflice", or "diflerent"
- Exact-match heading detectors silently skip clauses on scanned documents

## Not for this skill when
- The document is a born-digital PDF, use the embedded text layer instead of OCR
- The font has no ligatures and the miss comes from actual content differences
- The classifier fails on clean text, that is a model problem, not a glyph problem
- Non-Latin scripts where ligature behavior is different

## Variant phrasings
- fi ligature OCR error breaking clause matching
- tesseract misreads the fi glyph in scanned contracts
- OCR turned "liability" into "fiability" and the classifier missed it
- clause classifier missed the heading because of OCR glyph errors
- scanned indemnity clause not detected, OCR mangled the heading

## Why it happens
Many serif and legal fonts encode common pairs like fi, fl, and ff as a single glyph for typesetting. When the page is scanned, the OCR engine sees one connected shape instead of two letters and maps it to whatever its training data suggests, often a private-use codepoint or the wrong letter pair. Your pipeline then compares that mangled string against a clean heading list with exact equality, which fails on the first wrong character. The text was always there, the matcher was just too strict to see it.

## Edge cases
- Words that legitimately contain ligature pairs, like "sufficient", survive normalization unchanged, so the map is safe to apply globally.
- Some OCR engines have a preserve-ligatures option, check your engine's docs before adding your own pass, you may just need to flip a setting.
- Fuzzy matching can over-match short headings, keep a per-heading minimum length and a higher threshold for headings under 15 characters.
- If the engine emits the mangled letters rather than the ligature codepoint, step 2 does nothing and the fuzzy matcher in step 3 carries the fix, that is why both layers exist.
- Hand the classifier the normalized text but keep the original span offsets if you need to cite the source page, normalization shifts character positions.

## Provenance

Resolved from the public thread: https://vectle.com/posts/pst_8k7jKikVMW-z2lK2IOHzNw
