# Earnings transcript timestamps are broken after parsing

## TL;DR
Transcript timestamps break in parsing because every provider formats them differently, brackets versus parentheses versus bare minutes, and a parser written for one format mangles the rest. Normalize by detecting the actual timestamp pattern in the raw text first, then converting everything to seconds since call start. Validate the result by checking that timestamps increase monotonically through the call.

## The error
```text
(wrong data, no exception)
parsed transcript timestamps: [None, None, "0:00", "garbage"] / times out of order after parse
```

## When this helps
- parsed transcripts have missing or garbled timestamps
- aligning transcript quotes with call audio positions
- building a Q and A index by timestamp
- validating a new transcript parser

## When it doesn't
- the transcript has no timestamps at all; some providers omit them
- you need word-level alignment; turn-level timestamps cannot give that
- timestamps exist but speaker labels are wrong; that is a separate parse problem

## Works with
python 3.8+ with re. Transcript formats vary by provider (wires, IR sites, licensed APIs).

## Steps
### 1. Inspect the raw timestamp format before writing any parser
```bash
head -30 transcript_raw.txt
```
Expected: The actual timestamp style in the file. Look for the pattern: minutes and seconds with some wrapper, and whether every speaker turn carries one.

### 2. Detect the pattern programmatically
```python
import re
raw = open("transcript_raw.txt").read()
patterns = [r"\d{1,3}:\d{2}:\d{2}", r"\d{1,3}:\d{2}"]
for p in patterns:
    hits = re.findall(p, raw)
    print(p, "hits:", len(hits), "sample:", hits[:3])
```
Expected: Hit counts per pattern. The pattern with hits near the speaker-turn count is the real timestamp format; the others are durations inside the prose.

### 3. Normalize every timestamp to seconds
```python
import re
def to_seconds(ts):
    parts = [int(x) for x in ts.split(":")]
    while len(parts) != 3:
        parts = [0] + parts
    h, m, s = parts
    return h * 3600 + m * 60 + s
raw = open("transcript_raw.txt").read()
stamps = re.findall(r"\d{1,2}:\d{2}(?::\d{2})?", raw)
secs = [to_seconds(t) for t in stamps]
print("first:", secs[0], "last:", secs[-1], "count:", len(secs))
```
Expected: Seconds-since-start values. Normalizing once means every downstream consumer works in one unit.

### 4. Validate monotonic increase through the call
```python
import re
def to_seconds(ts):
    parts = [int(x) for x in ts.split(":")]
    while len(parts) != 3:
        parts = [0] + parts
    h, m, s = parts
    return h * 3600 + m * 60 + s
raw = open("transcript_raw.txt").read()
secs = [to_seconds(t) for t in re.findall(r"\d{1,2}:\d{2}(?::\d{2})?", raw)]
bad = [i for i in range(1, len(secs)) if min(secs[i], secs[i-1]) == secs[i] and secs[i] != secs[i-1]]
print("out-of-order positions:", bad[:10])
print("clean" if not bad else "timestamps need repair")
```
Expected: An empty bad list. Out-of-order hits are usually durations quoted in speech; filter to timestamps that precede speaker labels.

## Other ways people phrase this
### transcript timestamps none after parse
The parser's timestamp pattern did not match the provider's format. Detect first, then parse.

### earnings call timestamps out of order
Speech durations polluting the timestamp list. Anchor timestamps to speaker-turn boundaries.

### timestamp format mismatch transcript provider
Providers change formats without notice. The detect-then-normalize pattern survives format changes.

## Why it happens
There is no standard transcript timestamp format. Providers use different wrappers and granularities, and some stamp every turn while others stamp every few minutes. A parser with a hardcoded pattern silently drops or mangles the rest, and quoted durations in speech add false hits.

## Edge cases
- Calls longer than an hour need the hour component; a minutes-only pattern wraps at 60 and scrambles ordering.
- Prepared remarks sometimes carry no timestamps while Q and A does; handle the mixed case.
- Translated transcripts can reorder timestamp placement; validate per provider.
- Some providers timestamp by wall clock instead of call offset; convert with the call start time.

## Provenance

Resolved from the public thread: https://vectle.com/posts/pst_2YemXWWlDjet7KRMlTYlGA
