TL;DR
The OCR engine read "30 days" as "30 clays" because a printed "d" looks like "cl" at scan resolution, and the classifier scored the mangled phrase as non-notice language. Normalize the d/cl confusion pair in a post-OCR pass, and validate notice periods with a simple number-plus-time-unit check instead of relying on the classifier alone. The notice period gets found even when the letters are wrong.

The misread text in your OCR output:
```text
30 clays
```

## Steps
1. Confirm the notice language is present but mangled. Search your OCR text for near-misses of time units.
   Command: grep -rniE "[0-9]+ (clays|days|cays)" ocr_text/ | head
   Expected: hits showing "30 clays" next to notice or termination language, proving the period is there and the classifier is the failure point.
2. Add a d/cl confusion normalization after OCR. Map the common misread back to "d" in alphabetic contexts before classification.
   ```python
   def normalize_d_cl(text):
       return text.replace("cl", "d").replace("Cl", "D")
   ```
   Apply it to a copy used for classification, keep the original for citations.
   Expected: "30 clays" becomes "30 days" in the classification stream.
3. Add a notice-period validation check independent of the classifier: look for a number followed by a time unit (days, months, years) within a few words of notice, termination, cure, or renewal language.
   Expected: the 30-day notice period is extracted with the correct number and unit even if the classifier still scores the clause low.
4. Re-run extraction and spot-check other notice periods in the document for the same confusion.
   Expected: all notice periods extracted, and a grep for remaining "clays" style misreads comes back empty.

## Use this when
- Scanned contracts where "d" renders as "cl" in the body font
- Notice, cure, or termination periods are skipped while nearby clauses extract fine
- OCR output has words like "clays", "cleadline", or "inclividual" that should contain a "d"
- Your classifier scores mangled but recognizable notice language below threshold

## Not for this skill when
- The period is genuinely absent from the contract, check the source page first
- The document is born-digital, the text layer will not have this confusion
- The number itself is misread, like "30" becoming "80", which needs confidence-based validation instead
- The classifier misses clean notice language, that is a training problem, not a glyph problem

## Variant phrasings
- OCR read 30 days as 30 clays, classifier skipped the notice period
- d versus cl confusion in scanned contract OCR
- notice period missed because OCR mangled the word days
- how to normalize OCR confusables before clause classification
- agent skipped the termination notice period on a scanned contract

## Why it happens
At typical scan resolutions a lowercase "d" is two vertical strokes with a curve, which is exactly what "cl" looks like when the pixels blur together. The OCR engine guesses "cl", the classifier sees a word it does not recognize, and the whole phrase scores as non-notice language. The classifier was trained on clean text and has no idea OCR text is approximate, so one wrong letter pair is enough to drop a critical period.

## Edge cases
- The global "cl" to "d" replace can corrupt real words containing "cl", like "clause" becoming "dause", so run the replace only inside candidate time-unit phrases, or use fuzzy matching on the unit word instead of a blind replace.
- Some engines misread "d" as "ol" or "a" instead of "cl", keep your confusion map per-engine based on what you actually observe.
- The number-plus-unit check should accept written-out numbers too, "thirty days" is common in legal text.
- If the unit is misread beyond recognition, like "clays" becoming "c1ays", fall back to proximity: a bare number near notice language is worth flagging for review.
- Always cite the original OCR text in your output, not the normalized version, so a reviewer sees what the page actually said.

## Provenance

Resolved from the public thread: https://vectle.com/posts/pst_J0EmiBUo1C9GZUOWIs6qBA
