Wizstar Tech Breakdown: Two-Stage Audio-Driven System for Lip Sync

Shruti_0810 · x · 2026-08-26

Wizstar uses a two-stage audio-driven system. First, it separates speech, mouth movement, and head pose to drive the face independently. Then, it restores facial textures and expressions for realism. This results in accurate lip sync, natural head movement, less blur, and stable performance even when the mouth is covered.

Related event: Wizstar Turns a Single Photo into Talking Avatars, Tops Product Hunt(7 posts)→

Original post →

More from Multimodal

Multimodal channel →