Open-source LongCat-Avatar turns one photo plus audio into minutes of lip-synced talking video
Roger_M_Taylor · x · 2026-09-15
A viral post highlights LongCat-Avatar, an open-source and free project that generates minutes-long, lip-synced talking-head video from just a single photo and an audio clip — capabilities that previously required a camera, studio and editing. The poster claims it "humiliated the entire paid AI video industry," with the repo link shared in the comments.
More from Multimodal
- GPT-6 Astra + Hyper3D Rodin MCP Turns One Image Into a Full 3D Scene — heyshrutimishra · 2026-09-15
- RTX 5060 runs 12-sec H3 video with 2 ref images in 57 minutes — thatguyjames_uk · 2026-09-15
- 4060 Ti test: LTX 2.5 beats H3 on VRAM and speed, so why does everyone use H3? — ForesterAI · 2026-09-15
- MiniMax H3 hits 2x real-time video denoising with SGLang + VDN-H3 on 8x B200 — MiniMax_AI · 2026-09-15
- Ant Group's Realtime-Venus: full-duplex interaction with asynchronous model delegation — antgroup · 2026-09-15
- Odyssey Teases World Models Part 3 Release, Following Interactive Video Push — soleio · 2026-09-15