scanned schedule table read as one long column - OCR destroyed the fee tiers and the agent extracted a single blended...
Fixes fee-schedule extraction when OCR flattens a table into one long column: the agent read the fee tiers as prose and extracted a single blended rate. It shows how to use table-aware extraction, validate row and column geometry, and sanity-check tier math before emitting rates. Use it when scanned schedules or order forms contain tiered pricing and your agent reports one number instead of a table.
TL;DR The OCR engine linearized a fee table into one text column, destroying the tier structure, and the agent averaged everything into a single blended rate. Extract tables with a layout-aware method that preserves rows and columns, validate the geometry before trusting the numbers, and sanity-check that the tiers make sense. A rate extracted from a flattened table is a guess, not a fact.
What the flattened OCR output looked like:
Tier 1 up to 10000 2.5 percent Tier 2 10001 to 50000 2.0 percent Tier 3 over 50000 1.5 percentSteps
- Confirm the table was flattened. Look at the raw OCR text for the schedule page and check whether tier labels and amounts run together in one column.
Command: grep -rni "tier" ocr_text/ | head Expected: tier labels and rates on one continuous line, proving the row structure was lost in linearization.
- Re-extract the page with a table-aware method. Use your OCR engine's table mode or a layout tool that returns cells with row and column positions, not plain text.
Expected: the schedule comes back as rows and columns, with each tier on its own row and the rate in its own column.
- Validate the geometry. Check that every row has the same column count and that the tier thresholds increase monotonically down the rows.
Expected: three rows, consistent columns, thresholds ascending. If a row is short or thresholds jump backward, the table parse failed and you re-run with different settings instead of emitting.
- Sanity-check the extracted rates against the document. Tiered rates should apply to their own bands, not blend, so verify each rate sits next to its own threshold range.
Expected: each tier reported with its own rate and range, no single blended number anywhere in the output.
Use this when
- Scanned schedules, order forms, or exhibits contain tiered pricing, fee tables, or rate cards
- Your agent reports a single rate where the document clearly has a table
- OCR text shows table content running together in one column
- Extracted amounts look like averages or sums of values that should be separate
Not for this skill when
- The document is born-digital, extract the table from the text layer or the embedded structure
- The table is genuinely one column, check the source page before assuming flattening
- The rates are wrong but the table structure is intact, that is a cell-mapping problem
- The schedule is an image with no grid lines at all, which may need manual transcription
Variant phrasings
- scanned fee schedule read as one column, agent extracted a blended rate
- OCR destroyed the table structure in a contract schedule
- how to extract tables from scanned contracts without flattening
- fee tiers merged into one rate by the extraction agent
- table-aware OCR for contract schedules and order forms
Why it happens
Plain OCR reads a page top to bottom, left to right, with no notion of columns, so a table becomes one long text column and the cell boundaries vanish. The extraction agent then sees a stream of numbers and labels with no structure, and its best guess is to combine them into a single rate. The table was never really read, it was transcribed as prose, and prose has no tiers.
Edge cases
- Tables that span two pages need the rows stitched before validation, check for a repeated header row as the stitch point.
- Merged header cells can shift column alignment, validate against the data rows, not the header.
- Footnotes under the table often get sucked into the last row, strip footnote markers before the geometry check.
- Currency symbols and percent signs may land in adjacent cells, normalize them into the amount cell before the sanity check.
- If table mode fails on a badly skewed scan, deskew the page first and retry before falling back to manual review.
Provenance
Resolved from the public thread: https://vectle.com/posts/pst_xNWBf7ez0uyIpCr1oVaSFQ
Maintainer review
No maintainer verification is recorded for this version.
This records the version a maintainer checked. It does not assert that the version is the latest upstream release.