Production
Nobody benchmarks the one thing a product video is made of
PhyWorldBench scored twelve models across 12,600 videos on physics. The one benchmark for the text in the frame covers ten models and puts its leader at 0.37.
PhyWorldBench, published as a conference paper at ICLR 2026, generated 12,600 videos to test twelve text to video models against 1,050 physics prompts. The best of the twelve, Pika 2.0, passed 0.262 of them. Now search the paper's 74,631 characters of text for OCR, for legibility, for typography, for glyph: zero hits, for every one of those terms. The same search across VBench and VBench-2.0, the suites the field actually cites, returns zero as well. So the field has an expensive, careful, peer reviewed answer to whether a dropped ball falls correctly, and for whether the product name on screen is spelled right it has one human rubric and one opinion score. A marketing video is mostly text.
The leading score is 0.262, and it is an intersection of two judgements
PhyWorldBench scores each generated video twice, and both scores are either a zero or a one. Semantic adherence asks whether the video shows what the prompt asked for. Physical commonsense asks whether the physics holds. The headline figure is the share of videos that satisfy both, and Pika 2.0's 0.262 is that intersection. Its physical commonsense alone is 0.314. Its semantic adherence alone is 0.521. Quoting 0.262 as a physics accuracy is quoting the overlap of two different failures, and the overlap is always smaller than either one.
The suite also contains prompts that deliberately break physics, and every model declines from the fundamental categories to the composite ones to those. Pika 2.0 scores 0.043 on physical commonsense there and 0.011 on the joint measure, and five of the twelve models beat it on that joint column, Kling 1.6 highest at 0.042. The paper says why those prompts are in there: the gap between a model's physical and unphysical performance reveals whether it understands physics or merely replicates patterns from training data that mostly follows real world laws. The best model on ordinary physics lands mid table when the prompt asks for the impossible, which is what that sentence predicts.
One more thing about this figure, because it travels. The abstract dated 26 May 2026 says twelve models, including five open source and five proprietary. Five and five is ten. The body says five proprietary and seven open source, and the NVIDIA Cosmos Lab page for the same benchmark still says ten models, naming Hunyuan as the open source leader where the paper names Wanx-2.1.
Quoting 0.262 as a physics accuracy is quoting the overlap of two different failures.
Sixteen dimensions, then eighteen more, and none of them is text
VBench, submitted in November 2023 and presented at CVPR 2024, is the suite a model release cites when it wants to look measured. It has sixteen dimensions: subject consistency, background consistency, temporal flickering, motion smoothness, dynamic degree, aesthetic quality, imaging quality, object class, spatial relationship, appearance style and six more like them. Every axis is about the world in the shot. None is about the writing on it.
VBench-2.0 arrived in March 2025 arguing that the earlier axes mainly represent superficial faithfulness, whether a video looks visually convincing rather than whether it obeys real world principles. It adds five dimensions, human fidelity, controllability, creativity, physics and commonsense, split into eighteen fine grained capabilities: human anatomy, geometry multi-view consistency, state change thermotics, motion rationality, instance preservation, camera motion and twelve more. Not one of the eighteen concerns a character on screen either.
The claim is a grep, not a reading. Take the full extracted text of all three papers, 126,182 characters of VBench, 87,250 of VBench-2.0 and 74,631 of PhyWorldBench, and search case sensitively for OCR, legib, typograph, glyph and optical character. Every one of those five terms returns zero hits in every one of the three. Three of the most cited evaluation suites in the field, and the vocabulary for readable text never appears.
PhyWorldBench, ICLR 2026: twelve text to video models, 1,050 prompts each, 12,600 generated videos. The measure is the share of a model's videos that satisfy both semantic adherence and physical commonsense, so it is the intersection of two binary judgements rather than a physics score on its own. The axis is declared at 1.00, which is why the leader's bar looks the way it does.
The only benchmark for on-screen text puts its leader at 0.37, and disagrees with itself
T2VTextBench, posted on 8 May 2025, calls itself the first human evaluation benchmark dedicated to on-screen text fidelity and temporal consistency in text to video models. Ten systems, 73 prompts each, six categories, three student annotators scoring each output on a four point rubric where 0.25 means under half the text is correct and 0.5 means the majority is, at 50 to 80 per cent. Sora averages 0.37, the highest of the ten. So the leader sits below the majority correct band, and the paper states the general result plainly: every model averages below 0.4, indicating a consistent failure.
One of the six categories is App and Web UI Simulation, which is the product video case exactly, a screen with strings on it. Sora records 0.50 there, its best column of the six. Kling records 0.00 there and 0.00 in four other categories, for an overall average of 0.01. That is not a spread between good models and weaker ones. It is a spread between unreliable and unusable.
The paper also disagrees with itself. Table 2 is the results table. Table 3 is a cost table carrying its own average score column, and the two disagree for nine of the ten models: Pika at 0.44 against 0.36, LTX Video at 0.28 against 0.20, Sora at 0.36 against 0.37. Only Kling agrees, at 0.01 in both. The same model is called Pika 2.2 in one table and Pika 2.1 in the other. On Table 3 Pika leads; on Table 2 Sora does. The figures in this post come from Table 2. The published record on text in generated video is thin enough that one document's internal inconsistency decides who leads it.
The only other measurement is thinner still. AIGVE-60K covers 3,050 prompts across 20 fine grained task dimensions, 58,500 videos from 30 models and 120K mean opinion scores. Exactly one of the twenty dimensions is OCR, and the finding, in the authors' words, is that the 30 models show relatively poor alignment on the position, OCR, linguistic structure and complexity challenges, with scores clustering around 30 on a scale rescaled to 0 to 100. One human rubric, one opinion score. That is the whole evidence base.
Ask the same question about a still image and there is a metric, a leaderboard and a leader
OCRGenBench, whose current version is dated 20 March 2026, covers 33 OCR generative tasks under five text categories, document, handwriting, scene text, artistic text and layout-rich text, over 1,060 manually annotated samples, and defines a composite metric called OCRGenScore. Nano Banana Pro leads at 77.19, one of only two models above 70, most below 60 against a theoretical maximum of 100. The word video does not appear anywhere in the paper.
So the same question is answered to two decimal places for stills and not at all for video, and the reason is the instrument. A UCLA paper on video text preservation, from November 2025, states it directly: unlike image generation, where OCR tools can be reliably applied, text in video suffers from frame-wise inconsistency, distortion and low resolution, which makes standard metrics unreliable. That also explains why the failure is hard to catch by eye. A word legible in one frame is not the same thing as a word that stays the same word for ninety frames.
The vendor documentation is the only other place to look. Google DeepMind's Veo page sells real world physics, realism, prompt adherence and consistency. The Vertex AI best practices page for the same model mentions text in the frame once, under a heading called Avoid quotation marks, and the advice is how to stop it happening: use a colon after the speaker's action rather than quotation marks, to prevent the model from rendering text in the video. The prompt guide says nothing about on-screen text, and neither does the Gemini API reference. The one documented control over text in a Veo frame is a switch for turning it off.
The one documented control over text in a Veo frame is a switch for turning it off.
If nobody can measure it, do not ask a model to draw it
The product name. The price. The button label. The number in the case study. Those are the load bearing pixels of a marketing video, and every one of them is text. No tool that generates them can tell you how often they come out right, because nobody has built the instrument that would answer. The two studies that tried came back with a human rubric and an opinion score.
The fix is dull: do not generate the text, render it. Type set in HTML and CSS is the exact string that was typed, at the chosen size, in a typeface the company already owns, and it is identical in every frame of the shot because it is the same string in every frame. Nothing to score, because nothing to guess. Generate or capture the footage, where the benchmarks at least exist, and composite everything a viewer has to read.
This is where an earlier piece in this journal, on how much of the feed watches with the sound off, meets this one. If the argument has to be on screen because most of the audience never hears it, and the words on screen come out of a sampler whose error rate nobody has published, then two failure modes are stacked and only one is visible in the export. Watching the cut with the volume at zero catches the first. Nothing in the current literature catches the second.
Left, eight of the axes that VBench, VBench-2.0 and PhyWorldBench actually score. Right, the same frame a marketing video is made of, with its text marked. None of the three benchmarks has a dimension that reads any of it.
What to do about it
- Put every string that has to be readable into the script as a string, then composite it as type. Nothing a viewer has to read should come out of a sampler.
- When a text to video score is quoted at you, ask what it is a share of. PhyWorldBench's 0.262 is the intersection of semantic adherence and physical commonsense, not a physics grade.
- Before repeating any published text fidelity figure, check which table it came from. T2VTextBench's two tables disagree for nine of its ten models.
Questions people actually ask
is there a benchmark for text in ai generated video?
One, and it is small: T2VTextBench, posted on 8 May 2025, describes itself as the first human evaluation benchmark dedicated to on-screen text fidelity in text to video models, and it covers ten systems at 73 prompts each. The only other published measurement is a single OCR dimension inside the AIGVE-60K dataset, one of its 20 task dimensions. Neither reports character accuracy. VBench, with sixteen dimensions, and VBench-2.0, with eighteen fine grained ones, have no text dimension at all.
how accurately does ai video generate on-screen text?
The highest average in T2VTextBench is 0.37 out of 1, recorded by Sora across 73 prompts scored by three human annotators. On that paper's rubric 0.5 means the majority of the text is correct, at 50 to 80 per cent, so the leader sits below the majority correct band, and Kling averaged 0.01. In the App and Web UI category, the case closest to a product video, the best score any of the ten models reached was 0.50.
why can't you just run ocr on the frames of a generated video?
Because OCR is not reliable on video output. A UCLA paper on video text preservation, from November 2025, states that unlike image generation, where OCR tools can be reliably applied, text in video suffers from frame-wise inconsistency, distortion and low resolution, which makes standard metrics unreliable. A word can be legible in one frame, deformed in the next and gone in the third, and a per-frame character accuracy averages that into a number describing none of the three.
what does phyworldbench's 0.262 actually measure?
0.262 is the share of Pika 2.0's 1,050 generated videos that satisfied both of PhyWorldBench's binary judgements at once: semantic adherence, meaning the video shows what the prompt asked for, and physical commonsense, meaning the physics holds. Pika's physical commonsense alone is 0.314 and its semantic adherence alone is 0.521. It is the best result of the twelve models tested across 12,600 videos, and it is an intersection rather than a physics grade.
how should text be added to a marketing video?
Render it as type rather than asking a model to draw it, because the highest on-screen text score any of the ten models in T2VTextBench recorded was 0.37 out of 1. Type set in HTML and CSS is the exact string that was typed, at the chosen size, and it is identical in every frame of the shot. Generate or capture the footage, then composite the product name, the price and the interface labels over it.
Sources
- arXiv, PhyWorldBench: A Comprehensive Evaluation of Physical Realism in Text-to-Video Models
- NVIDIA Research Cosmos Lab, PhyWorldBench project page
- arXiv, T2VTextBench: A Human Evaluation Benchmark for Textual Control in Video Generation Models
- arXiv, VBench: Comprehensive Benchmark Suite for Video Generative Models
- arXiv, VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness
- arXiv, LOVE: Benchmarking and Evaluating Text-to-Video Generation and Video-to-Text Interpretation, the AIGVE-60K dataset
- arXiv, Video Text Preservation with Synthetic Text-Rich Videos
- arXiv, OCRGenBench: A Comprehensive Benchmark for Evaluating OCR Generative Capabilities
- Google Cloud, Best practices for Veo on Vertex AI
- Google DeepMind, Veo
- Google AI for Developers, Generate videos with Veo
- Google Cloud, Video generation prompt guide
Every figure on this page comes from one of these. Where two of them measure the same thing differently, the article says so rather than picking the flattering one.