earnings transcript timestamps broken after parsing
This skill fixes broken earnings transcript timestamps after parsing. Use it when timestamps come back missing, garbled, or out of order, or when validating a new transcript parser. It is not for transcripts without timestamps; the fix is detecting the provider's actual pattern, normalizing to seconds, and validating monotonic order.
Earnings transcript timestamps are broken after parsing
TL;DR
Transcript timestamps break in parsing because every provider formats them differently, brackets versus parentheses versus bare minutes, and a parser written for one format mangles the rest. Normalize by detecting the actual timestamp pattern in the raw text first, then converting everything to seconds since call start. Validate the result by checking that timestamps increase monotonically through the call.
The error
(wrong data, no exception)
parsed transcript timestamps: [None, None, "0:00", "garbage"] / times out of order after parseWhen this helps
- parsed transcripts have missing or garbled timestamps
- aligning transcript quotes with call audio positions
- building a Q and A index by timestamp
- validating a new transcript parser
When it doesn't
- the transcript has no timestamps at all; some providers omit them
- you need word-level alignment; turn-level timestamps cannot give that
- timestamps exist but speaker labels are wrong; that is a separate parse problem
Works with
python 3.8+ with re. Transcript formats vary by provider (wires, IR sites, licensed APIs).
Steps
1. Inspect the raw timestamp format before writing any parser
head -30 transcript_raw.txtExpected: The actual timestamp style in the file. Look for the pattern: minutes and seconds with some wrapper, and whether every speaker turn carries one.
2. Detect the pattern programmatically
import re
raw = open("transcript_raw.txt").read()
patterns = [r"\d{1,3}:\d{2}:\d{2}", r"\d{1,3}:\d{2}"]
for p in patterns:
hits = re.findall(p, raw)
print(p, "hits:", len(hits), "sample:", hits[:3])Expected: Hit counts per pattern. The pattern with hits near the speaker-turn count is the real timestamp format; the others are durations inside the prose.
3. Normalize every timestamp to seconds
import re
def to_seconds(ts):
parts = [int(x) for x in ts.split(":")]
while len(parts) != 3:
parts = [0] + parts
h, m, s = parts
return h * 3600 + m * 60 + s
raw = open("transcript_raw.txt").read()
stamps = re.findall(r"\d{1,2}:\d{2}(?::\d{2})?", raw)
secs = [to_seconds(t) for t in stamps]
print("first:", secs[0], "last:", secs[-1], "count:", len(secs))Expected: Seconds-since-start values. Normalizing once means every downstream consumer works in one unit.
4. Validate monotonic increase through the call
import re
def to_seconds(ts):
parts = [int(x) for x in ts.split(":")]
while len(parts) != 3:
parts = [0] + parts
h, m, s = parts
return h * 3600 + m * 60 + s
raw = open("transcript_raw.txt").read()
secs = [to_seconds(t) for t in re.findall(r"\d{1,2}:\d{2}(?::\d{2})?", raw)]
bad = [i for i in range(1, len(secs)) if min(secs[i], secs[i-1]) == secs[i] and secs[i] != secs[i-1]]
print("out-of-order positions:", bad[:10])
print("clean" if not bad else "timestamps need repair")Expected: An empty bad list. Out-of-order hits are usually durations quoted in speech; filter to timestamps that precede speaker labels.
Other ways people phrase this
transcript timestamps none after parse
The parser's timestamp pattern did not match the provider's format. Detect first, then parse.
earnings call timestamps out of order
Speech durations polluting the timestamp list. Anchor timestamps to speaker-turn boundaries.
timestamp format mismatch transcript provider
Providers change formats without notice. The detect-then-normalize pattern survives format changes.
Why it happens
There is no standard transcript timestamp format. Providers use different wrappers and granularities, and some stamp every turn while others stamp every few minutes. A parser with a hardcoded pattern silently drops or mangles the rest, and quoted durations in speech add false hits.
Edge cases
- Calls longer than an hour need the hour component; a minutes-only pattern wraps at 60 and scrambles ordering.
- Prepared remarks sometimes carry no timestamps while Q and A does; handle the mixed case.
- Translated transcripts can reorder timestamp placement; validate per provider.
- Some providers timestamp by wall clock instead of call offset; convert with the call start time.
Provenance
Resolved from the public thread: https://vectle.com/posts/pst_2YemXWWlDjet7KRMlTYlGA