Meta launches Muse Realtime Avatar, turning voice streams into live talking characters
alexandr_wang · x · 2026-09-24
- Meta AI Research introduced Muse Realtime Avatar, which turns Muse Realtime Voice's speech token stream into expressive, interactive avatars in real time: portraits show subtle expressions, full-body illustrations gesture and shift posture, and even animals or objects stay expressive while keeping a consistent identity across conversational turns.
- Technically, voice and embodiment share one streaming pipeline: an audio-driven Diffusion Transformer conditioned on speech tokens, reference media, and a rolling window of recent video latents generates video in short causal chunks, bounding computation so generation can run for the whole conversation.
- Everything generated is watermarked as AI with no added latency, forming the foundation for Meta's realtime embodied AI products. Meta also ran head-to-head tests against Runway Characters and HeyGen LiveAvatar, saying people preferred Muse.
Related event: Meta Launches Muse Realtime Avatar, Beats HeyGen and Runway in Blind Tests(3 posts)→
More from Multimodal
- Creator open-sources brushstroke animation workflow built on Claude Opus 5.5 — alejandroll10 · 2026-09-24
- Meta's 3-year AI arc: from Threads to the Muse video model — minchoi · 2026-09-24
- Meta Connect showcases team's realtime interaction work — EdwardSun0909 · 2026-09-24
- 14-minute fully AI-generated series "Bride Of The Atom" Episode 1 released — geekycheekypixels · 2026-09-24
- DS 4.1 Flash drives Blender 3D animation, hinting multimodal training is essential for visual art — bookwormengr · 2026-09-24
- Muse adds voice generation; Scale AI CEO quips 'time to Pixar larp' — alexandr_wang · 2026-09-24