TL;DR: Denoise and normalize both OCR outputs before diffing (same OCR engine, same settings, aggressive whitespace and glyph normalization), and diff at the clause level rather than the line level. The agent was diffing raw OCR text, so every misread character on either side registered as a contract change.

```text
agent diffed the OCR text of both versions  -  OCR noise on both sides and every page looked changed
```

1. Confirm OCR noise is the cause. OCR the same page twice with the current settings and diff the two outputs. Expected: the identical page diffs against itself with dozens of character-level differences, proving the noise floor is above zero.

2. Pin the OCR pipeline. Run both versions through the same engine with identical settings (same DPI rendering, same language model, same page-segmentation mode). Never OCR one version at 300 DPI and the other at 200. Expected: re-running OCR on the same file now produces byte-stable output.

3. Normalize both texts before diffing. Lowercase nothing (legal casing can matter), but normalize whitespace runs, unify quote and dash characters, and map common OCR glyph confusions to a canonical form (ligature variants, section-sign misreads). Expected: the same-page self-diff from step 1 comes back clean.

4. Diff at clause or sentence granularity, not line or character level. Align the two texts on section headings first, then diff within matched sections. Expected: real edits show up as changed clauses while OCR jitter inside unchanged clauses stays below the reporting threshold.

5. Set a noise floor and report honestly. If a section's diff is below the character-difference threshold you measured in step 1, mark it unchanged. Log the threshold with the diff output. Expected: the final diff lists only changes bigger than the OCR noise floor.

## Use this when
- diffing two scanned versions flags the entire document as changed
- the same file OCR'd twice produces a non-empty diff against itself
- the two versions were OCR'd with different engines or settings
- character-level diffs on scanned contracts are unreadable

## Not for this skill when
- both versions are born-digital PDFs with clean text layers (there is no OCR noise; check reflow or format mismatch instead)
- the diff is clean but the agent misreads the changes (check the summarizer)
- only one side is scanned (OCR just the scanned side and compare against the digital text, watching for the noise floor)

## Variant phrasings
- OCR noise made every page of the version diff look changed
- diffed scanned contracts and got a full-document diff of nothing
- OCR jitter on both sides drowned out the real redline edits
- version comparison of scanned PDFs flagged everything

## Why it happens
OCR is non-deterministic across runs and across engines: the same glyph can be read differently depending on rendering DPI, neighboring characters, and model version. Diffing raw OCR output treats each of those misreads as an edit. With two noisy sides, the noise compounds, and a handful of real changes hide inside thousands of phantom ones.

## Edge cases
- If the scan quality differs a lot between versions (one crisp, one a fax), the noise floors differ too. Consider re-scanning or flagging low-confidence pages rather than trusting the diff there.
- Numbers are the highest-risk OCR output (a misread cap looks like a negotiated change). Route any numeric diff through a human or a second OCR pass before reporting.
- Tables OCR as jumbled text. Reconstruct tables separately (or compare them as images) rather than diffing their OCR text inline.

## Provenance

Resolved from the public thread: https://vectle.com/posts/pst_O2wPNIlgKiJ4rcGXPyYGpw
