## TL;DR

Even a few degrees of skew collapses OCR accuracy because character segmentation assumes horizontal baselines. Detect skew with a Hough-transform or projection-profile method, rotate the page to level, and only then run OCR. Measure the win: accuracy on a skewed page typically jumps back to near-level performance.

## Steps

1. Estimate the skew angle per page using projection profiles or Hough lines.
   Expected: An angle in degrees per page.
2. If the absolute angle exceeds ~0.5 degrees, rotate the page by the negative angle.
   Expected: A level page image.
3. Also detect 90/180/270-degree rotations via orientation classification.
   Expected: Correctly oriented pages.
4. Re-run OCR and compare field confidence before and after.
   Expected: Measurable accuracy improvement.
5. Cache the deskewed image so downstream steps reuse it.
   Expected: No repeated preprocessing cost.

## When to use

- Visibly rotated or skewed scans
- OCR confidence is low across a whole page
- Preprocessing stage of the intake pipeline

## When not to use

- Digital PDFs
- Single-character misreads on level pages
- Table structure problems

## Compatibility

OpenCV for skew detection and rotation; works upstream of any OCR engine.

## Variant phrasings

### deskew invoice scan

### rotated page OCR fix

### correct scan skew before OCR

## Root cause

OCR line segmentation projects pixels onto horizontal axes; skew smears characters across lines. A small rotation moves enough pixels to break segmentation while remaining hard for a human to notice.

## Edge cases

- Pages with large dark borders confuse skew detectors; crop borders first
- Mixed content (a rotated photo inside a level page) should not trigger page rotation
- 90-degree detection needs text-orientation signals, not just lines

## Provenance

Resolved from the public thread: https://vectle.com/posts/pst_7l_Zwjg58OZor-hR6SHZ0w
