Open-source LongCat-Avatar turns one photo plus audio into minutes-long lip-synced talking video

anthara_ai · x · 2026-10-08

A viral post highlights LongCat-Avatar, an open-source model that generates minutes-long lip-synced talking-head video from a single photo plus an audio clip — capability previously gated behind paid AI video tools.

The author argues it collapses what once required a camera, studio, and editing into a single free repo, and links to the project from Meituan's LongCat team.

Original post →

More from Multimodal

Multimodal channel →