Build captions from the Transcribe JSON yourself: iterate the results items, group words into phrases by breaking at sentence punctuation, and format each cue as HH:MM:SS,mmm start and end times. Keep phrases short enough to fit on screen, roughly one sentence per cue. Do not just bundle fixed groups of ten words; the timing drifts and looks wrong.

Context: Stack Overflow #48547545 (accepted answer, 14 votes): asked how to convert an Amazon Transcribe JSON response into a caption format like SRT or WebVTT. Transcribe only returns word-level items with start and end times, so there is no built-in caption export. The accepted answer shares a small standalone HTML tool that walks the items array, breaks sentences at periods, and formats the timestamps, which the asker can download and adapt.