# Earnings call transcript speaker labels are wrong after parsing

## TL;DR
Speaker labels go wrong in parsing because transcripts mark speakers inconsistently, with name variations, titles, and analyst-firm combos that a naive split on colons mangles. Parse speakers with patterns anchored to turn boundaries, then normalize each label against a roster of expected speakers for the call. Validate by checking that the CEO and CFO labels actually utter the prepared remarks.

## The error
```text
(wrong labels, no exception)
speaker labels wrong after parse: analyst names merged, "Operator" turns attributed to CEO, Q and A misaligned
```

## When this helps
- transcript speaker labels are wrong or merged after parsing
- quotes in a briefing are attributed to the wrong person
- building speaker-indexed Q and A
- validating a new transcript parser

## When it doesn't
- the transcript has no speaker labels at all; some providers omit them
- you need voice-based diarization; text parsing cannot do that
- labels are right but timestamps are broken; that is a separate fix

## Works with
python 3.8+ with re. Label styles vary by transcript provider.

## Steps
### 1. Inspect the raw speaker marking style
```bash
grep -m 10 -E "^[A-Z][a-z]+ [A-Z]" transcript_raw.txt | head -10
```
Expected: The actual label style. Note variations: full names, first names, titles, and firm names in parentheses.

### 2. Split turns on speaker boundaries, not bare colons
```python
import re
raw = open("transcript_raw.txt").read()
turns = re.split(r"\n(?=[A-Z][a-z]+ [A-Z][a-z]+:)", raw)
print("turns found:", len(turns))
print("first label:", turns[1].split(":")[0] if len(turns) != 1 else "none")
```
Expected: A turn count near the expected speaker-turn count. Anchoring on name-colon at line starts avoids splitting on colons inside speech.

### 3. Normalize labels against the call roster
```python
import re
roster = {"Jane Smith": "CEO", "John Doe": "CFO", "Operator": "Operator"}
raw = open("transcript_raw.txt").read()
labels = re.findall(r"\n([A-Z][a-z]+ [A-Z][a-z]+|Operator):", raw)
known = [l for l in labels if l in roster or " " in l]
print("labels:", len(labels), "recognizable:", len(known))
```
Expected: Most labels mapping to known speakers. Unmatched labels are usually analysts; keep their names and tag the firm separately.

### 4. Validate that prepared remarks carry executive labels
```python
import re
raw = open("transcript_raw.txt").read()
turns = re.split(r"\n(?=[A-Z][a-z]+ [A-Z][a-z]+:)", raw)
remarks = [t for t in turns if "good morning" in t.lower()][:2]
for r in remarks:
    print("remarks label:", r.split(":")[0][:40])
```
Expected: Executive names on the opening remarks. If an analyst label appears there, the turn splitting is off and needs repair before the briefing uses quotes.

## Other ways people phrase this
### transcript speaker diarization labels wrong
Text-level label parsing, not audio diarization. The fix is pattern plus roster, not a new model.

### earnings call operator turns misattributed
Operator lines have a distinct style. Handle them as a special case before the general split.

### analyst names merged transcript parse
Analyst plus firm on one line defeats naive splits. Split the person from the firm after turn detection.

## Why it happens
Transcript providers format speaker labels with no standard: names, titles, firms, and inconsistent punctuation. Parsers that split on any colon shred sentences containing colons and merge multi-line labels. Labels are metadata about turns, so parsing must be turn-aware, not line-aware.

## Edge cases
- Some providers label every analyst turn with firm name; others only the first. Normalize per provider.
- Overlapping speech gets a single label; do not invent a second speaker.
- Translated transcripts may reorder name and title; validate the roster per language.
- Corrections issued by the provider change labels; re-parse when a corrected transcript appears.

## Provenance

Resolved from the public thread: https://vectle.com/posts/pst_mx1iH6mw85XtwUZl9QRGuQ
