Local face swap works, but lips don't: Reddit compares Wav2Lip vs MuseTalk for lip sync
krypto_nighte · reddit · 2026-10-09
A Reddit user running the H3 character-swap LoRA workflow on an RTX 3060 12GB with 32GB RAM gets clean swaps at 14 min per 832x640 clip — but the swapped face keeps the source video's mouth movements when new dialogue is added.
Their comparisons so far:
- Wav2Lip: leaves a visible box around the lips at this resolution
- MuseTalk: better, but flickers on side turns
- Hosted lipsync services work but they'd rather keep everything in the local graph
They're asking what others chain after the swap, and whether to lipsync before or after upscaling — a practical pitfall reference for anyone building local video face-swap/dubbing workflows.
More from Multimodal
- 740M-param cross-modal embeddings for video/audio/code run on consumer edge silicon — clmt · 2026-10-09
- Alibaba open-sources Qwen-Image-2.1-Turbo, generating 2K images in just 8 denoising steps — Alibaba_Qwen · 2026-10-09
- GenIA: Rendering-Guided Test-Time Alignment Turns SAM3D into SOTA Image-to-3D — JonathonLuiten · 2026-10-09
- Recreating a Ben 10 Scene with MiniMax H3: 10 Retries and Manual Editing — nUclear_nOva89 · 2026-10-09
- Smudged Eraser Ghost Layers prompt: make subjects emerge from half-wiped chalk residue — LudovicCreator · 2026-10-09
- VNCCS PoseStudio LoRA redraws any character in a 3D mannequin pose on Qwen-Image — linoy_tsaban · 2026-10-09