Meta details Muse's unified streaming architecture: one speech-token stream drives voice and avatar
AIatMeta · x · 2026-09-24
Meta explains how Muse Realtime Voice and Muse Realtime Avatar form a single streaming system: Voice generates speech tokens encoding content and prosody, which Avatar consumes to generate synchronized video. A fixed-length history serves as motion context for subsequent chunks, keeping computation bounded over arbitrary conversation lengths while keeping voice, lip motion and expressions in sync.
Related event: Meta Unveils Muse Realtime Avatar: Sub-second Realtime Digital Humans(18 posts)→
More from Multimodal
- Claude Opus 5.5 generates a full launch video — animation, music, voiceover — in 20 minutes — cedric_chee · 2026-09-26
- Feeding YuE 2 an empty lyric field yields a song full of gibberish vocals — SteveLittleFish · 2026-09-26
- "Opus 5.5" Rumored Release Draws Rave First Impressions and Cynicism — chaumian · 2026-09-26
- Reddit user shares striking clip: 'Video models are getting good' — we_are_mammals · 2026-09-26
- Gemma plays Snake straight from pixels via VLM gateway, under 240ms p99 at ~$0.00007/image — spillai · 2026-09-26
- Opus 5.5 directs a sci-fi short on the Arecibo message via Krea MCP and Hyperframes — angrypenguinPNG · 2026-09-26