Meta details Muse Realtime Voice: VQ tokens drive parallel audio and video decoders
alex_conneau · x · 2026-09-24
Muse Realtime Voice continuously generates VQ tokens: a speech decoder reconstructs audio while an audio2video decoder produces video chunks for realtime avatars. The author calls diffusion-based videogen for avatars "brute force" but argues the same heavily optimized architecture can extend to richer voice-driven interactive worlds — realtime video adds expression, embodiment and context on top of voice.
Related event: Meta Unveils Muse Realtime Avatar: Sub-second Realtime Digital Humans(18 posts)→
More from Models
- GPT-6-Astra-Max Tops BALROG Game Benchmark at 68.3%, NetHack Depth Hits 13.2 — burny_tech · 2026-09-26
- Claude Opus proves surprisingly good at both making IK rigs and animating with them — andrew_n_carr · 2026-09-26
- Two years since o1-preview: reasoning tokens went from novelty to frontier standard — ArtificialAnlys · 2026-09-26
- Open-source JevBench adds $49/$99 paid priority benchmark runs with 48-hour turnaround — airesearch12 · 2026-09-26
- Ling Tiny 3.0 on a 2017 laptop: 8B model hits 10 tok/s, hints at edge AI future — netherreddit · 2026-09-26
- Watch Claude Opus 5.5 generate 100 matplotlib plots in one go — goodside · 2026-09-26