termination notice address missed - OCR merged the footer text into the clause and the agent's page segmentation...
Fixes a contract agent that missed the termination notice address because OCR merged footer text into the clause and page segmentation dropped the region. Use when addresses, dates, or notice details vanish from extraction on scanned contracts. Not for clauses missing from the source document, digital PDFs with a clean text layer, or redline-diff false positives.
TL;DR
Footers and margin text get fused into clause text by OCR, and then the page-segmentation step discards the whole region as non-content. The fix is to segment the page image into header, body, and footer zones before OCR, extract each zone separately, and run clause extraction on the body while pulling notice addresses from the footer zone explicitly.
The query
termination notice address missed - OCR merged the footer text into the clause and the agent's page segmentation dropped itUse this when
- Addresses, dates, or notice details vanish from extraction on scanned contracts.
- The missing content sits in a footer, header, or margin in the page image.
- Re-running OCR with layout preservation shows the text the pipeline dropped.
Not for
- Clauses missing from the source document itself.
- Digital PDFs with a clean text layer (no OCR or segmentation involved).
- Redline-diff false positives between versions.
Steps
Step 1: Confirm the address is in a dropped zone
tesseract page-14.png stdout --psm 6 | grep -i -A 2 "notice"Expected output: the notice address appears in the raw OCR output when layout is preserved. If it appears here but not in the pipeline's extraction, the segmentation step dropped it, and the fix belongs in zoning, not in OCR quality.
Step 2: Add a zoning pass before text extraction
# split each page image into zones by vertical position
zones = {
"header": page.crop(top=0.00, bottom=0.08),
"body": page.crop(top=0.08, bottom=0.92),
"footer": page.crop(top=0.92, bottom=1.00),
}Expected output: three separate images per page. The footer band that carried the notice address is now its own extraction input instead of noise fused into the clause above it.
Step 3: Extract zones separately and keep them labeled
for zone in header body footer; do
tesseract page-14-$zone.png page-14-$zone.txt
doneExpected output: three text files with zone labels. Footer text never flows into clause paragraphs again, because the zones never mix before extraction.
Step 4: Run clause extraction on the body, address extraction on the footer
python extract_clauses.py --input page-14-body.txt
python extract_notice_address.py --input page-14-footer.txtExpected output: clauses come from the body zone, the notice address comes from the footer zone. Each extractor reads the zone where its content actually lives.
Step 5: Add a completeness check for notice details
Add an extraction check: a termination clause without a notice address is incomplete and must be flagged for review.
Expected output: the pipeline now treats a missing notice address as a failed extraction, not a successful one with less data. The next dropped footer surfaces as a flag, not a silent gap.
Step 6: Apply zoning per page across the whole document
python zone_all_pages.py contract.pdf --out zones/Expected output: every page zoned the same way. Footers repeat on every page, so the zoning must run per page. Deduplicate footer content after extraction, not before, or the one page whose footer differs gets lost.
Variant phrasings
agent skipped the auto-renewal date, coffee stain on the scanned page
Same dropped-content class, different cause (image damage instead of zoning). Flag low-confidence regions for human review.
scanned schedule table read as one long column
Same layout-loss family: the table structure died in OCR. Extract tables with a table-aware tool before the text pipeline.
OCR merged the header letterhead into the first clause
Same zoning fix, top of the page instead of the bottom. The header zone exists for exactly this.
Why it happens
OCR is line-oriented, and page-segmentation models trained on articles and books treat footers as noise to discard. Legal footers are not noise: they carry notice addresses, execution dates, and page-level operative text. The pipeline fused two zones into one text stream, then a segmentation step trained on the wrong document genre threw away the part that mattered. The content was never missing. The pipeline's model of the page was wrong.
Edge cases
- Some contracts put the notice address in the body on the signature page and in the footer everywhere else. Extract from both zones and reconcile, preferring the signature-page instance.
- Zone boundaries by fixed percentage fail on pages with unusual layouts (full-page exhibits, landscape pages). Detect zones by content (repeated text across pages signals a footer) as a fallback.
- Do not deduplicate footer text before extraction. The one footer that differs (an amended notice address on the last page) is usually the operative one.
- Multi-column body layouts need column-aware zoning too, or the body extraction interleaves columns the same way the footer fused into the clause.
- Keep the zone-labeled OCR output alongside the final extraction. When an address is disputed, the auditor needs to see which zone it came from.
Provenance
Resolved from the public thread: https://vectle.com/posts/pst0mnRGPPKr1Hr4gBe7-HkA
Maintainer review
No maintainer verification is recorded for this version.
This records the version a maintainer checked. It does not assert that the version is the latest upstream release.