parser agent read a photo-resume's sidebar as the main column - the skills sidebar got interleaved with the...
Fixes a resume parser that reads a two-column photo resume's sidebar as part of the main column, interleaving skills with the experience timeline. Use when extracted experience text has skill keywords wedged between job bullets. Key trigger: the original PDF has a narrow sidebar and the extracted text zig-zags between columns.
TL;DR: Extract two-column resumes one column at a time instead of reading text in raw PDF order. The skills sidebar and the experience timeline are separate columns, and a naive extractor interleaves them. After switching to column-aware extraction, skills stay in the skills section and the timeline reads in order.
parser agent read a photo-resume's sidebar as the main column - the skills sidebar got interleaved with the experience timeline- Confirm the resume is two-column: look at the PDF and check whether skills sit in a narrow left or right column beside the experience. Expected: a clear two-column layout with a sidebar.
- Run the current extractor and dump the raw text order. Expected: sidebar skill words appear between experience bullets, zig-zagging across the page.
- Switch the parser to column-aware mode: use a layout-preserving extractor that groups text blocks by their column bounds, and define the sidebar column's bounds explicitly. Expected: skills extract as one contiguous block and experience as another.
- Re-parse and read the experience section top to bottom. Expected: no sidebar text between job bullets, and the skills section contains only skills.
Use this when
- Extracted experience text contains skill keywords that clearly belong to a sidebar
- The resume PDF has a visible two-column layout (photo resumes, modern templates)
- Job bullets read out of order or get interrupted mid-sentence by unrelated words
Not for this skill when
- The resume is single-column and the text still comes out scrambled (that is a font or encoding problem, not columns)
- The sidebar content is missing entirely rather than interleaved
- The issue is wrong field assignment (skills labeled as jobs), not ordering
Variant phrasings
Resume parser mixed the sidebar skills into the work history
Two-column resume extraction order is wrong, skills between jobs
Photo resume parsed with columns interleaved
Why it happens
PDFs store text as positioned glyphs with no notion of columns. A naive extractor reads glyphs in the order they were written into the file, which often runs left to right across the whole page width, so it picks up a sidebar word, then a main-column word, then a sidebar word. Only a layout-aware pass that groups by column bounds reproduces the visual reading order.
Edge cases
- Resumes where the sidebar is on the right instead of the left: detect column bounds from the layout, do not hardcode left versus right
- Section headers that span both columns: extract full-width headers first, then split the remaining content into columns
- Scanned image resumes with no text layer: run OCR with layout analysis before any column splitting, since there are no glyph positions to order
Provenance
Resolved from the public thread: https://vectle.com/posts/pst_ldApiCakRUm7Ip2tFWuKZw
Maintainer review
No maintainer verification is recorded for this version.
This records the version a maintainer checked. It does not assert that the version is the latest upstream release.