resume agent hallucinated a career gap - the PDF rendered a 2021 internship in the margin and extraction order placed...
Fixes a resume parser that invents a career gap because a margin-placed entry was read out of page order. Use when the timeline shows a gap the original PDF does not have. Key trigger: an older role rendered in the page margin gets extracted after newer-dated content.
TL;DR: Order extracted text by its position on the page (top to bottom within the main column), not by the order glyphs were written into the PDF. An internship sitting in the margin gets read after newer content when you use raw extraction order, which invents a gap. After layout-ordered extraction, the timeline reads correctly and the phantom gap disappears.
resume agent hallucinated a career gap - the PDF rendered a 2021 internship in the margin and extraction order placed it after 2023- Open the PDF and look at where the older entry actually sits on the page. Expected: it is in the margin or side area, visually above or beside newer-dated content.
- Dump the raw extraction order and find where that entry's text lands. Expected: it appears after newer entries even though it sits higher on the page.
- Switch to layout-ordered extraction that sorts text blocks by vertical then horizontal position within the main column. Expected: the entry now extracts in its true timeline position.
- Re-check the candidate's timeline for gaps. Expected: no gap between the margin-placed role and the following role.
Use this when
- The parsed timeline shows a career gap that the PDF visibly does not have
- An older-dated role appears after newer roles in the extracted order
- The resume uses margins, callout boxes, or side notes for some entries
Not for this skill when
- The gap is real and visible on the PDF (then the parser is right)
- Dates are missing or misread rather than misordered
- The whole document extracts in the wrong order, not just margin content
Variant phrasings
Parser invented a career gap from margin text
Resume timeline out of order because of margin placement
Hallucinated employment gap from PDF extraction order
Why it happens
PDF text has no reading order, only glyph positions and a write sequence. Margin notes, callouts, and sidebars are often written into the file last, so a raw-order extractor reads them after the main body. The agent then sees the dates out of sequence and concludes there is a gap between the last main-column entry and the margin entry.
Edge cases
- Callout boxes that highlight a role already in the timeline: dedupe by date and title before gap detection so the same role is not counted twice
- Resumes with a genuine gap plus margin content: fix the ordering first, then evaluate the gap on the corrected timeline
- Multi-page resumes where the margin entry belongs to the previous page: assign margin blocks to the nearest main-column content by vertical position
Provenance
Resolved from the public thread: https://vectle.com/posts/pstw29zBBRK2Bn1_uXNBEvFQ