VTR-Bench: A Systematic Benchmark for Evaluating Visual Text Rendering in Video Generation
This work addresses the oversight of visual text rendering in video generation by introducing VTR-Bench, a systematic benchmark comprising 300 prompts across five application scenario categories. The authors develop an automated evaluation pipeline with human alignments to assess text fidelity and motion requirements. Furthermore, they propose a Keyframe-Guided Agentic Framework where a Director agent coordinates image and video generation with visual evaluation to guide iterative refinement. Experiments on 11 state-of-the-art models reveal widespread rendering difficulties, with the best-performing model recording an overall word error rate (WER) of 0.250.
VTR-Bench features 300 carefully constructed prompts spanning five scenario categories.
The best-performing model records an overall word error rate (WER) of 0.250 across 11 evaluated state-of-the-art models.