resume agent pulled the reference letter header as the candidate's current job - the recommender's company name got...
Stops a resume parser from mistaking reference-letter headers for employment history by segmenting the document into resume vs attachments before parsing. Use when parsed employment history contains a recommender's company or other non-resume content. Not for overlapping-role double counting, sidebar layout misreads, or cover-letter header misclassification.
TL;DR
Attached reference letters look like employment history to a naive parser: company name, title, dates, all in the same text stream. The fix is document segmentation. Split the PDF into resume vs attachments before parsing, using letter markers like salutations and signature blocks. Only the resume zone feeds employment extraction. Everything else is context, not history.
The query
resume agent pulled the reference letter header as the candidate's current job - the recommender's company name got parsed as employment historyUse this when
- Parsed employment history contains a recommender's company or other non-resume content.
- The resume PDF bundles reference letters, cover letters, or transcripts after the resume.
- A candidate's "current job" matches a letterhead, not their actual employment.
Not for
- Overlapping contract roles double-counted as separate jobs (that is a date-merge problem).
- Photo-resume sidebars interleaved with the timeline (that is a layout problem).
- The target role read from a cover-letter header (adjacent, but cover letters need their own rule).
Steps
Step 1: Confirm the bogus job matches the letter header
grep -i -B 3 -A 3 "[recommender company]" parsed-candidate.jsonExpected output: the parsed "job" entry whose company, title, and dates mirror the reference letter's header. If the entry has no bullet points and its dates match the letter date, segmentation is the problem.
Step 2: Detect letter-like pages before parsing
LETTER_MARKERS = [
"to whom it may concern",
"dear hiring manager",
"sincerely,",
"recommendation",
]
# any page containing a salutation plus a signature block is a letter, not a resume pageExpected output: a classifier that labels each page as resume or letter. Letters have salutations, body prose, and signature blocks. Resumes have section headers, bullets, and date ranges. The shapes barely overlap once you look.
Step 3: Parse employment only from the resume zone
python segment_documents.py application.pdf --out resume.pdf --out letters.pdf
python parse_resume.py resume.pdf --out candidate.jsonExpected output: candidate.json contains employment history drawn only from the resume pages. The letters land in a separate file, attached as context for the recruiter, never fed to the employment extractor.
Step 4: Add a sanity check on parsed jobs
python -c "import json; d=json.load(open('candidate.json')); print([j for j in d['jobs'] if not j.get('bullets')])"Expected output: any job with no bullets and no clear date range gets flagged for review. Real jobs have content. Letter headers do not.
Step 5: Re-parse and verify the recommender's company is gone
python parse_resume.py resume.pdf | grep -i "[recommender company]" || echo "clean"Expected output: "clean." The recommender's company no longer appears in employment history. It remains visible in the attached letters where it belongs.
Step 6: Log segmentation decisions for audit
Log the segmentation decision in segmentation-log.txt: application.pdf - pages 1-2 resume, pages 3-4 letter (Acme Corp).
Expected output: an audit trail of what was classified as what. When a recruiter asks why a job is missing, the log answers in one line.
Variant phrasings
parser credited a master's degree from the resume template's example text
Same contamination class, different source: template placeholder text parsed as real data. Strip known template markers before extraction.
resume agent assigned freelance client names as employers
Same flat-text-stream problem: gig listings under each contract got merged as jobs. Segment gig entries from employment entries by structure.
agent extracted the wrong email, picked the footer disclaimer address
Same zone confusion: the parser read the whole document instead of the header zone. Restrict contact extraction to the resume header.
Why it happens
Parsers treat the PDF as one flat text stream, and attachments inherit the resume's structure by proximity. A reference letter has everything an employment entry has (organization name, a person's title, dates), so a pattern matcher cannot tell them apart without document-level context. The parser was never wrong about the pattern. It was wrong about which document it was reading.
Edge cases
- Some candidates put the reference letter first and the resume second. Segment by content markers, never by page order.
- A candidate who works at the same company as their recommender creates a genuine ambiguity. Keep both, but tag the letter-derived one as unverified rather than merging silently.
- Scanned letters need OCR before segmentation. Run layout-aware OCR first, then the marker classifier on the text.
- Do not discard the letters. Recruiters want them. The fix is routing them to the right consumer, not deleting them.
- Multi-candidate PDFs (a recruiter's merged file) need per-candidate segmentation first. Page-level letter detection assumes one candidate per document.
Provenance
Resolved from the public thread: https://vectle.com/posts/pst_EG-eGEKeM42UOGIN-bkskg
Maintainer review
No maintainer verification is recorded for this version.
This records the version a maintainer checked. It does not assert that the version is the latest upstream release.