Meta Unveils Muse Realtime Avatar: Speech-Driven Real-Time Avatars
alex_conneau · x · 2026-09-25
Meta AI introduces Muse Realtime Avatar, extending Muse Realtime Voice into expressive real-time interactive avatars.
- Capabilities: Conditioned on reference media, any character — photographic portraits, full-body illustrations, animals, objects — comes alive with facial, hand, and full-body motion, staying coherent across conversational turns
- Architecture: Voice and Avatar share a single streaming speech token (VQ) stream for synchronized speech, lip motion, and expression; the avatar is an audio-driven Diffusion Transformer generating short causal chunks with a rolling window of recent latents, bounding computation so generation can run as long as the conversation
- Led by Amaury Sabran over nearly two years (started at Waveforms AI), solving lip-sync, distillation, long-rollout consistency, efficiency, hand gestures, and any-style support; former Meta voice lead Alexandre Conneau publicly praised the work
Related event: Meta Unveils Muse Realtime Avatar: Sub-second Realtime Digital Humans(18 posts)→
More from Multimodal
- Claude Opus 5.5 generates a full launch video — animation, music, voiceover — in 20 minutes — cedric_chee · 2026-09-26
- Feeding YuE 2 an empty lyric field yields a song full of gibberish vocals — SteveLittleFish · 2026-09-26
- "Opus 5.5" Rumored Release Draws Rave First Impressions and Cynicism — chaumian · 2026-09-26
- Reddit user shares striking clip: 'Video models are getting good' — we_are_mammals · 2026-09-26
- Gemma plays Snake straight from pixels via VLM gateway, under 240ms p99 at ~$0.00007/image — spillai · 2026-09-26
- Opus 5.5 directs a sci-fi short on the Arecibo message via Krea MCP and Hyperframes — angrypenguinPNG · 2026-09-26