Align Then Reason: A Multimodal Lip-Sync Judge for Dubbing
The authors introduce Align Then Reason (ATR), a multilingual lip-sync judge that evaluates dubbing quality using only silent video and candidate text. ATR establishes a monotonic alignment between frame-level lip representations and the phonetic units of the candidate line, providing both local evidence and a calibrated global alignment score to an LLM reasoner. Across a seven-language benchmark and downstream tasks, ATR demonstrates substantial improvements in mean AUC and outperforms existing lip-reading baselines on dub-line reranking and script-to-clip assignment.
On a seven-language benchmark, our method improves mean AUC over the corresponding Qwen3.5 SFT baselines by 59.4%, 50.2%, and 50.8% with 2B, 4B, and 9B reasoners, respectively.
The gains generalize across LLM families, reaching mean AUC improvements of 45.9% and 46.6% over the best baseline for LLaMA-3.1-8B and Mistral-7B, respectively.
On dub-line reranking, ATR-9B outperforms the best lip-reading baseline by 52.0%, while on script-to-clip assignment, ATR-9B improves over the best lip-reading baseline by 17.7%.