Alibaba's Wan 3.0 video model shows stable lip-sync across cuts and languages

AIwithGhotai · x · 2026-08-25

Alibaba's Wan 3.0 video generation model is now available on Magnific. Tests indicate its lip-sync capabilities hold up surprisingly well, even through scene cuts and mid-scene language switches—a common failure point for other video models.

The model supports generating 30 seconds of video with sound in a single API call, eliminating the need to choose between text-to-video and image-to-video modes.

Related event: Alibaba's Wan 3.0 Lands on Magnific with Superior Lip Sync and Consistency(7 posts)→

Original post →

More from Multimodal

Multimodal channel →