## TL;DR
Parsers flatten a resume to plain text and lose the section that tells a human "hobbies" apart from "work history." Tag sections before extracting, only credit skills found in experience or education, and never generate a duration without an explicit date range. Provenance on every claim keeps inflated experience out of the candidate record.

## The query

```text
resume parser agent hallucinated 3 years of React experience - the candidate's PDF had 'React' in the hobbies section and the agent counted it as professional experience
```

## Use this when

- A parser credits skills or years the candidate never earned at work
- Resume mentions in hobbies, summaries, or sidebars leak into experience
- Overlapping contract roles get summed into impossible totals

## Not for

- OCR quality problems on scanned resumes
- ATS API errors or sync issues
- Interview scheduling bugs

## Steps

### 1. Tag sections before extracting anything

Split the resume into labeled sections (experience, education, skills, hobbies, summary) with heading detection or layout analysis before any skill or date extraction runs.

Expected output: every text span carries a section label, and hobbies text is marked as hobbies.

### 2. Only credit skills from experience and education

Mentions in hobbies, summary, or template boilerplate get flagged as unverified interest, never counted as professional experience.

Expected output: 'React' in the hobbies section no longer adds years to the candidate's record.

### 3. Require date ranges before crediting duration

A years-of-experience claim needs an explicit date range inside the same job entry. A bare keyword match never generates a duration on its own.

Expected output: no duration attached to a skill without dates to back it up.

### 4. Merge overlapping ranges instead of summing

Two roles both dated 2021-2023 running in parallel count once. Sort ranges, merge overlaps, then total.

Expected output: de-duplicated totals that cannot exceed the candidate's actual working years.

### 5. Emit provenance per claim

Store the skill, the source section, the date range, and a confidence score with every extracted fact.

Expected output: downstream consumers (ranking, screening agents) see exactly where each claim came from and can discount shaky ones.

## Variant phrasings

### resume parser counted a hobby as work experience

Same pipeline. Step 2 is the fix: section-scoped crediting. A hobby mention can still surface as a soft signal, but it must never become years of experience.

### agent inflated the candidate's years of experience

Check steps 3 and 4 first. Most inflation comes from duration-without-dates plus naive summing of overlapping roles.

## Why it happens

Most parser pipelines convert the PDF to a flat text stream and run keyword extraction, which destroys the section structure a human reads effortlessly. Layout order also interleaves sidebars into the timeline, so a skills sidebar bleeds into job entries. Without section awareness, every mention counts the same.

## Edge cases

- Resumes with no headings: fall back to positional heuristics (first block is contact info, dated blocks are experience) and lower the confidence.
- Functional resumes that list skills with no dates: report the skills as unverified, not as zero years.
- Photo resumes with sidebar layouts: run layout analysis first. Interleaved extraction order is the top source of cross-section contamination.

## Provenance

Resolved from the public thread: https://vectle.com/posts/pst_RyPXglscADZXx7VZOp2nYg
