Meta details Muse's unified streaming architecture: one speech-token stream drives voice and avatar

AIatMeta · x · 2026-09-24

Meta explains how Muse Realtime Voice and Muse Realtime Avatar form a single streaming system: Voice generates speech tokens encoding content and prosody, which Avatar consumes to generate synchronized video. A fixed-length history serves as motion context for subsequent chunks, keeping computation bounded over arbitrary conversation lengths while keeping voice, lip motion and expressions in sync.

Related event: Meta Unveils Muse Realtime Avatar: Sub-second Realtime Digital Humans(18 posts)→

Original post →

More from Multimodal

Multimodal channel →