TL;DR
A coffee stain turned "2027" into "202?" and the agent silently skipped the auto-renewal date instead of flagging the garbage character. Never let a date containing non-digit characters pass extraction. Flag low-confidence or non-digit characters in dates, re-OCR the damaged region after cleaning the image, and route anything still uncertain to a human. A missing date with a loud flag beats a skipped date every time.

The corrupted text in your OCR output:
```text
202?
```

## Steps
1. Find every date-like field your pipeline extracted and check for non-digit characters or low per-character confidence.
   Command: grep -rEn "[0-9]{3}[^0-9]" ocr_text/ | head
   Expected: hits like "202?" showing exactly which dates are corrupted instead of clean.
2. Add a validation gate: any extracted date containing a character that is not a digit or a known separator fails extraction and gets flagged for review, never emitted as a fact.
   Expected: the auto-renewal date comes out as "uncertain, needs review" rather than being silently skipped or emitted as "202?".
3. Re-OCR the damaged region. Crop the page region around the date, run denoising and contrast enhancement, and OCR the crop at a higher resolution than the full page.
   Expected: the cleaned crop reads "2027" with high per-character confidence, or stays uncertain and keeps the review flag.
4. If the region is still unreadable after cleanup, escalate to a human with the cropped image attached.
   Expected: a reviewer confirms the date from the crop, and your extraction log records the date as human-verified, not OCR-verified.

## Use this when
- Scanned pages have stains, folds, stamps, or highlighter over key fields
- OCR output contains "?" or other garbage characters inside dates, amounts, or clause numbers
- Your agent silently skips fields it cannot read instead of flagging them
- Key dates like renewal, termination, or effective dates come from damaged scans

## Not for this skill when
- The date is cleanly read but wrong, that is a different problem, check your date parsing logic
- The document is born-digital, the text layer will not have stain artifacts
- The uncertainty is about which date applies, like three conflicting renewal dates, that is a reconciliation problem
- Your OCR engine already exposes confidence scores you are ignoring, wire those in first

## Variant phrasings
- OCR turned 2027 into 202? and the agent skipped the renewal date
- coffee stain on scanned contract, OCR misread the date
- damaged scan corrupted a key date, agent never flagged it
- how to handle low confidence OCR characters in contract dates
- agent silently dropped a date it could not read

## Why it happens
OCR engines output their best guess per character and mark uncertainty with substitute characters like "?" or with low confidence scores. Most extraction pipelines read the text and ignore both signals, so a corrupted date either gets emitted as garbage or, worse, the field fails a parse and the agent quietly moves on without the date. The stain is only half the problem; the other half is a pipeline that treats OCR text as ground truth instead of as a guess with error bars.

## Edge cases
- Stamps and signatures overlap dates often, crop generously around the date region so cleanup has context to work with.
- Handwritten corrections near a date can be the actual operative value, flag those for human review rather than trusting either the print or the handwriting.
- Deny-listing "?" is not enough, some engines substitute other characters, so also reject dates where any character confidence falls below your threshold.
- A date that reads cleanly but disagrees with other dates in the document is a reconciliation issue, flag the conflict rather than picking one silently.
- Keep the original scan region linked to the flagged date so the reviewer never has to hunt for the page.

## Provenance

Resolved from the public thread: https://vectle.com/posts/pst_ZtveLewqc1mOYLFxFsDyuw
