## TL;DR

Attached reference letters look like employment history to a naive parser: company name, title, dates, all in the same text stream. The fix is document segmentation. Split the PDF into resume vs attachments before parsing, using letter markers like salutations and signature blocks. Only the resume zone feeds employment extraction. Everything else is context, not history.

## The query

```text
resume agent pulled the reference letter header as the candidate's current job - the recommender's company name got parsed as employment history
```

## Use this when

- Parsed employment history contains a recommender's company or other non-resume content.
- The resume PDF bundles reference letters, cover letters, or transcripts after the resume.
- A candidate's "current job" matches a letterhead, not their actual employment.

## Not for

- Overlapping contract roles double-counted as separate jobs (that is a date-merge problem).
- Photo-resume sidebars interleaved with the timeline (that is a layout problem).
- The target role read from a cover-letter header (adjacent, but cover letters need their own rule).

## Steps

### Step 1: Confirm the bogus job matches the letter header

```bash
grep -i -B 3 -A 3 "[recommender company]" parsed-candidate.json
```

Expected output: the parsed "job" entry whose company, title, and dates mirror the reference letter's header. If the entry has no bullet points and its dates match the letter date, segmentation is the problem.

### Step 2: Detect letter-like pages before parsing

```python
LETTER_MARKERS = [
    "to whom it may concern",
    "dear hiring manager",
    "sincerely,",
    "recommendation",
]
# any page containing a salutation plus a signature block is a letter, not a resume page
```

Expected output: a classifier that labels each page as resume or letter. Letters have salutations, body prose, and signature blocks. Resumes have section headers, bullets, and date ranges. The shapes barely overlap once you look.

### Step 3: Parse employment only from the resume zone

```bash
python segment_documents.py application.pdf --out resume.pdf --out letters.pdf
python parse_resume.py resume.pdf --out candidate.json
```

Expected output: `candidate.json` contains employment history drawn only from the resume pages. The letters land in a separate file, attached as context for the recruiter, never fed to the employment extractor.

### Step 4: Add a sanity check on parsed jobs

```bash
python -c "import json; d=json.load(open('candidate.json')); print([j for j in d['jobs'] if not j.get('bullets')])"
```

Expected output: any job with no bullets and no clear date range gets flagged for review. Real jobs have content. Letter headers do not.

### Step 5: Re-parse and verify the recommender's company is gone

```bash
python parse_resume.py resume.pdf | grep -i "[recommender company]" || echo "clean"
```

Expected output: "clean." The recommender's company no longer appears in employment history. It remains visible in the attached letters where it belongs.

### Step 6: Log segmentation decisions for audit

Log the segmentation decision in segmentation-log.txt: application.pdf - pages 1-2 resume, pages 3-4 letter (Acme Corp).

Expected output: an audit trail of what was classified as what. When a recruiter asks why a job is missing, the log answers in one line.

## Variant phrasings

### parser credited a master's degree from the resume template's example text
Same contamination class, different source: template placeholder text parsed as real data. Strip known template markers before extraction.

### resume agent assigned freelance client names as employers
Same flat-text-stream problem: gig listings under each contract got merged as jobs. Segment gig entries from employment entries by structure.

### agent extracted the wrong email, picked the footer disclaimer address
Same zone confusion: the parser read the whole document instead of the header zone. Restrict contact extraction to the resume header.

## Why it happens

Parsers treat the PDF as one flat text stream, and attachments inherit the resume's structure by proximity. A reference letter has everything an employment entry has (organization name, a person's title, dates), so a pattern matcher cannot tell them apart without document-level context. The parser was never wrong about the pattern. It was wrong about which document it was reading.

## Edge cases

- Some candidates put the reference letter first and the resume second. Segment by content markers, never by page order.
- A candidate who works at the same company as their recommender creates a genuine ambiguity. Keep both, but tag the letter-derived one as unverified rather than merging silently.
- Scanned letters need OCR before segmentation. Run layout-aware OCR first, then the marker classifier on the text.
- Do not discard the letters. Recruiters want them. The fix is routing them to the right consumer, not deleting them.
- Multi-candidate PDFs (a recruiter's merged file) need per-candidate segmentation first. Page-level letter detection assumes one candidate per document.

## Provenance

Resolved from the public thread: https://vectle.com/posts/pst_EG-eGEKeM42UOGIN-bkskg
