Netflix's Align Then Reason lip-sync judge boosts mean AUC by up to 59%
netflix · hf · 2026-10-02
Netflix introduces Align Then Reason (ATR), a multilingual lip-sync judge for dubbing QC that checks whether a candidate text line matches a speaker's articulation in content and timing using only silent video and text.
- Prior lip-reading and video-language models are insensitive to temporal errors; ATR first builds monotonic alignment between frame-level lip representations and the candidate line's phonetic units, then lets an LLM reason jointly over per-phoneme local evidence and a calibrated global alignment score;
- On a seven-language benchmark it improves mean AUC over Qwen3.5 SFT baselines by 59.4%/50.2%/50.8% with 2B/4B/9B reasoners; gains of 46% transfer across LLaMA-3.1-8B and Mistral-7B and to three unseen MuAViC languages;
- On real dubbing tasks, ATR-9B beats the best lip-reading baseline by 52.0% on dub-line reranking and 17.7% on script-to-clip assignment.
More from Multimodal
- Image editing demo with Nano Banana — tkasasagi · 2026-10-02
- Creator: Opus 5.5 edits videos well, but forcing AI to clip without real need yields garbage — AlchainHust · 2026-10-02
- Reddit user's VEC concept car AI video shows startlingly realistic motion — Vashukanni · 2026-10-02
- Fable 5.5 rumored: webpage morphs its art style to match passing images — lxfater · 2026-10-02
- Oxford VGG Unveils SynCity 3000, Generating Globally Coherent Scene-Scale 3D Worlds — rsasaki0109 · 2026-10-02
- Omni-Embed-Mini: A 0.9B Embedder Adds Five Modalities Without Touching Text Weights — _reachsumit · 2026-10-02