Meta details Muse Realtime Voice: VQ tokens drive parallel audio and video decoders

alex_conneau · x · 2026-09-24

Muse Realtime Voice continuously generates VQ tokens: a speech decoder reconstructs audio while an audio2video decoder produces video chunks for realtime avatars. The author calls diffusion-based videogen for avatars "brute force" but argues the same heavily optimized architecture can extend to richer voice-driven interactive worlds — realtime video adds expression, embodiment and context on top of voice.

Related event: Meta Unveils Muse Realtime Avatar: Sub-second Realtime Digital Humans(18 posts)→

Original post →

More from Models

Models channel →