Production
A caption that cannot disagree with the script
AI caption tools run recognition over finished audio and guess the words and the times. Forced alignment takes the text as an input and can only decide when.
WhisperX was published at Interspeech 2023 with a table nobody quotes back at it. On the AMI test set, sixteen meeting recordings, WhisperX places 60.3 per cent of the reference words inside a 200 millisecond window with the string spelled correctly. Whisper's own timestamps manage 52.1. Read the metric before you read the number: a word only scores if the time is close and the text is an exact match, so the 39.7 per cent that failed is a mixture of words that were mistimed and words that were misheard, and the figure cannot tell you which. Every consumer caption tool is scored this way, because each guesses at both. A pipeline holding an approved script is not.
The transcript is an argument to the aligner, not a result of it
TorchAudio's tutorial states the shape in three sentences. The waveform is passed to an acoustic model, which produces the sequence of probability distribution of tokens. The transcript is passed to a tokenizer, which converts it to a sequence of tokens. The aligner takes the results from the acoustic model and the tokenizer and generates timestamps for each token. Two inputs. One output, and the output is time.
The API reference is blunter. The aligner's job is to align a label sequence to an emission, the text arrives as a target sequence, and what comes back is a label for each time step in the alignment path. The words are not in the return type. Merge the token spans and the tutorial prints a line like quante, score 1.00, from 3.055 to 3.678 seconds: a word you supplied, with two numbers on it.
The Montreal Forced Aligner defines the operation with the transcription first: take an orthographic transcription of an audio file and generate a time aligned version, using a pronunciation dictionary to look up phones for words. That dictionary is the commercially interesting part, because it is how a product name gets told how it is said. One condition comes with all of this: the tutorial says the process expects that the input transcript is already normalized, so numerals, acronyms and units become writing decisions.
The words are not in the aligner's return type, so there is no path by which it hands back a different one.
The best published word timing is 60.3 per cent, and it cannot say which half failed
Whisper's utterance timestamps are prone to inaccuracies, the WhisperX authors write, and word level timestamps are not available out of the box. A footnote to their Table 2 says Whisper's word times come from dynamic time warping over the attention scores of the decoded tokens. A third architecture, with the time read off attention weights rather than off a word boundary.
Their metric is the honest part of the paper. A true positive is a predicted word segment overlapping a ground truth segment within a 200 millisecond collar, where both words are an exact string match. On AMI they report 84.1 precision and 60.3 recall, against 78.9 and 52.1 for Whisper alone. Their own conclusion: Whisper alone underperforms on word segmentation precision and recall on both corpora, falling short even of wav2vec2.0.
A caption you can read is a claim about the word. One that lands on the frame is a claim about the time. The 60.3 is both claims holding at once, so it cannot say which one failed.
Stops from the TorchAudio 2.8 forced alignment tutorial and its aligner API reference. The words are an input at stop one and never appear in the output: the only thing this machine decides is when each of them happens.
A classical aligner beats both recognisers, and the study had to throw words away to measure it
Rousso, Cohen, Keshet and Chodroff published a direct comparison on 27 June 2024, running the Montreal Forced Aligner against WhisperX and MMS on two manually aligned corpora, TIMIT and Buckeye. The classical aligner won at every tolerance on both. Per cent of TIMIT word boundaries correct inside 10 milliseconds: 41.6 for the Montreal Forced Aligner, 22.4 for WhisperX, 18.6 for MMS. At 50 milliseconds, 89.4, 82.4 and 75.7.
Comparisons were conducted only on words correctly recognized by WhisperX and MMS. TIMIT's reference transcript holds 39,834 words; WhisperX got 37,685 right, so 2,149 could not be scored, and MMS got 29,057, so 10,777 could not. On Buckeye's 285,347 words, MMS lost 26,158. Every timing figure you have read about a recogniser was computed on the subset it heard.
The two papers put opposite numbers on the same system and both are right. WhisperX reports 60.3 per cent recall at a 200 millisecond collar and looks respectable. Rousso and colleagues report 22.4 per cent of TIMIT boundaries inside 10 milliseconds and it looks poor. A 200 millisecond collar is a caption tolerance and 10 milliseconds is a phonetics tolerance, and caption work lives at the first until a word has to trigger a cut.
Their Table 3 carries the pair worth planning around: WhisperX on Buckeye, median shift 30.1 milliseconds, mean 11,685.3. The authors explain the mean rather than let it stand as a verdict: Buckeye utterances run to several minutes, and a long input lets alignment drift. Their own 500 millisecond cutoff brings it to 36.4; the Montreal Forced Aligner's uncut mean on that corpus is 976.5. A median of 30 under a mean in the thousands means nearly every word is right and one is seconds out.
The words a founder cares about are the ones a recogniser misses most
A headline word error rate averages over text that is mostly ordinary English. Split the named entities out and one system reports two accuracies. On 25 recorded animal behaviour lectures, 17 hours of audio, a 2026 paper puts Whisper large-v3 at a 32.3 per cent word error rate on named entities and 7.0 per cent on every other word. Their own revision step pulls the entity rate to 22.7 and leaves the 7.0 untouched, so their best result is still three times worse on names than on ordinary words.
That is lecture audio, not a studio read, so take it as the shape of the gap rather than a forecast. A recogniser is worst at proper nouns, and a marketing video is almost entirely proper nouns.
W3C's accessibility guidance states the consequence flatly. Automatically generated captions do not meet user needs or accessibility requirements, unless they are confirmed to be fully accurate. Neither error in their worked example looks like one: broil on high for 4 to 5 minutes comes out as 45 minutes, and you should not preheat the oven comes out as you should know to preheat the oven. OpenAI's documented remedy is a prompt to improve recognition of names, acronyms, formatting or recording specific vocabulary, which makes a guess more likely to land. It does not make the word a given.
Nobody's caption rule is 99 per cent, and for prerecorded video there is no number at all
The 99 per cent figure circulates as the FCC caption accuracy requirement. Open FCC 14-12, adopted 20 February 2014, and as an accuracy figure it appears twice, both times as arithmetic inside a voluntary best practices list for real time captioning vendors: once in the discussion, once in the codified appendix. Seven thousand words, minus 70 errors, divided by 7,000, is 99 per cent. The order's own regulatory analysis says the captioning quality standards are qualitative rather than quantitative.
For prerecorded programming the Commission declined to put a number on it at all. Because of the greater opportunity for reviewing and editing captions on pre-recorded programming, it writes, such programming is capable of achieving full compliance except for de minimis errors. The record was not short of numbers: commenters proposed 100 per cent, 99.9, 99.8, and Media Captioning Services exactly 99 for post production work. None was adopted, and the reasoning is this post's argument: a recorded video can be reviewed before it airs.
The rule with the force of law names the specific failure. 47 CFR 79.1 requires captioning to match the spoken words in the order spoken, without substituting words for proper names and places, and without paraphrasing. Then read its definitions: video programming does not include advertisements of five minutes' duration or less. So the one American rule that forbids a phonetic product name exempts almost every marketing video ever cut, and your pricing page answers to WCAG 2.2 at Level A instead.
The 99 per cent everyone quotes is a worked division sum in a best practices list for live captioners.
Alignment adds under ten per cent to the run, so the approval is the expensive part
The WhisperX authors measured their alignment stage: the overhead is minimal, under ten per cent in speed. ElevenLabs sells forced alignment as a documented endpoint at the same rate as its speech to text API, across the 29 languages of its multilingual v2 models. What separates the two calls is whether you hold text somebody approved.
YouTube has offered both for years, on one help page, without naming the difference. Auto-sync asks you to enter the words in the video or upload a transcript file, then says transcript text is automatically synchronized to your video. One box asks for your text and returns times. The other asks for nothing and returns both.
Alignment has one real failure mode, and it is what makes the approval load bearing rather than ceremonial: the transcript has to be the text that was actually read. YouTube requires it be in the same language spoken in the video. ElevenLabs documents that diarised text will produce unexpected results. The CrisperWhisper authors name the structural cost, that discrepancies between model transcripts can further degrade timestamp precision. They also measure what noise does to a two model stack: on 200 synthesised samples with hand annotated timestamps, WhisperX word segmentation F1 falls from 76.7 to 59.0 at a signal to noise ratio of 1:5, a decline they put down to the wav2vec2.0 model it aligns against.
The left path is the shape of a transcription call, where word timing arrives as a parameter on the endpoint that also invents the words. The right path is the shape TorchAudio, the Montreal Forced Aligner and the ElevenLabs alignment endpoint all document. Its documented failure is a transcript that does not match the audio.
What to do about it
- Approve the script as text before the voice is recorded, and make that file the caption source. Anything that re-derives the words from the audio has discarded the only copy nobody had to guess at.
- Ask a caption tool one question: does it take my text as an input? If the answer is that it transcribes and you correct it afterwards, you are proofreading a guess on every re-render.
- Proofread the finished captions for proper nouns first, because that is where a recogniser's error rate runs several times its rate on ordinary words.
Questions people actually ask
what is forced alignment, and how is it different from transcription?
Forced alignment takes a transcript you already have plus the audio and returns two numbers for every word, a start time and an end time. TorchAudio describes the machinery in one line: the waveform goes to an acoustic model, the transcript goes to a tokenizer, and the aligner takes both results and generates timestamps for each token. Transcription runs the other way and produces the words as well as the times. The practical difference is that an aligner cannot spell a product name wrong, because the spelling was an input.
how accurate are automatic captions?
The best published word level figure is 60.3 per cent, from the WhisperX paper's test on 16 AMI meeting recordings, where a word scores only if its predicted segment overlaps the reference within 200 milliseconds and the two strings match exactly. Whisper's own timestamps score 52.1 per cent on the same test. Because the metric demands the right word and the right time together, it cannot say which of the two failed on the remaining 39.7 per cent.
why do AI captions get product names and brand names wrong?
Named entities are where recognition is weakest, and the gap is measured. On 25 recorded animal behaviour lectures, Whisper large-v3 shows a 32.3 per cent word error rate on named entities against 7.0 per cent on every other word of the same test set. The same authors cite 28.9 per cent on person names and 19.6 per cent on organisations for ConEC, a benchmark built from earnings calls. Neither corpus is a studio read, so both describe the shape of the gap rather than forecasting a voiceover, and recognisers fail plausibly rather than obviously: W3C's own example has broil on high for 4 to 5 minutes come out as 45 minutes.
is 99 per cent caption accuracy an FCC requirement?
No. The 99 per cent figure sits in FCC 14-12, adopted 20 February 2014, as arithmetic inside a voluntary best practices list for real time captioning vendors: 7,000 words minus 70 errors, divided by 7,000. Commenters did propose numbers for prerecorded programming, including 100 per cent, 99.9, 99.8 and 99, and none was adopted. The Commission expects full compliance except for de minimis errors, and the order's own regulatory analysis calls the standards qualitative rather than quantitative. A marketing video on a company website falls outside 47 CFR 79.1 altogether, and the obligation there comes from WCAG 2.2 at Level A.
can you upload your own transcript to youtube instead of using auto captions?
Yes, through Auto-sync, one of the two routes on YouTube's captions help page and forced alignment under another name. The page asks you to enter the words in the video or upload a transcript file, and then states that transcript text is automatically synchronized to your video. Two conditions come with it: the transcript must be in the same language spoken in the video, and transcripts are not recommended for recordings over an hour long or with poor audio quality.
Sources
- PyTorch, Forced alignment for multilingual data, TorchAudio 2.8 tutorial
- PyTorch, torchaudio.functional.forced_align API reference
- Bain, Huh, Han and Zisserman, WhisperX: Time-Accurate Speech Transcription of Long-Form Audio
- Rousso, Cohen, Keshet and Chodroff, Tradition or Innovation: A Comparison of Modern ASR Methods for Forced Alignment
- Wagner, Thallinger and Zusag, CrisperWhisper: Accurate Timestamps on Verbatim Speech Transcriptions
- Montreal Forced Aligner 3.x, User Guide
- ElevenLabs Documentation, Forced Alignment
- OpenAI, Speech to text guide
- YouTube Help, Add subtitles and captions
- FCC 14-12, Closed Captioning of Video Programming, Report and Order, adopted 20 February 2014
- 47 CFR 79.1, Closed captioning of televised video programming, Cornell LII
- W3C WAI, Captions and Subtitles
- Improving Speech Recognition of Named Entities in Classroom Speech with LLM Revision and Phonetic-Semantic Context
Every figure on this page comes from one of these. Where two of them measure the same thing differently, the article says so rather than picking the flattering one.