Stress-testing lip sync models with nine reproducible failure cases, from beards to fast speech
NewPhoneWhotiz · reddit · 2026-10-01
A developer is building a failure-case benchmark for lip sync/video workflows, cataloging nine recurring problem cases: hand over mouth, profile/near-profile views, tiny faces in wide shots, heavy beards, laughing, head turns while speaking, multiple speakers, fast speech, and hard cuts mid-word.
He plans to run the same source clip + audio pairing through both local/open workflows and hosted tools (like sync.so), logging versions, settings and source clips with a contact sheet for reproducibility—rather than judging from cherry-picked talking-head demos. So far every tool struggles with at least a few cases; he's soliciting missing categories.
More from Multimodal
- This 30-second vlog-style promo ad was made from a single image and one prompt — umesh_ai · 2026-10-01
- One prompt, under $20, one night: how Gradio distilled a 4-step text-to-image model — Gradio · 2026-10-01
- Distilled 260M text-to-image model hits <190ms on a T4, generating per keystroke — Gradio · 2026-10-01
- Prompt share: a fill-in template for minimalist editorial character illustrations — azed_ai · 2026-10-01
- Midjourney SREF Style Code 2041416583 Shared for Reuse — tisch_eins · 2026-10-01
- Supertonic TTS repo archived: development ended Sep 9, weights moved to archive namespace — JafarNajafov · 2026-10-01