Meta Introduces Muse Realtime Avatar: Speech-Driven Live Avatars at ~870ms Latency
alex_conneau · x · 2026-09-24
Meta AI Research unveils Muse Realtime Avatar, extending Muse Realtime Voice into expressive, interactive avatars with 870ms response latency, coming to Muse and Muse Charm.
Key technical points:
- Voice and avatar share a single speech-token (VQ) stream: the audio decoder turns tokens into speech while the avatar model consumes the same stream, keeping voice, lip motion, and expressions synchronized.
- The avatar is an audio-driven Diffusion Transformer conditioned on speech tokens, reference media, and a rolling window of recent video latents, generating video in short causal chunks that carry appearance and mannerisms forward indefinitely.
- Beyond talking heads: full-body illustrations gesture as they speak, and animals or everyday objects become expressive without losing their distinctive character.
Related event: Meta Unveils Muse Realtime Avatar: Sub-second Realtime Digital Humans(18 posts)→
More from Models
- GPT-6-Astra-Max Tops BALROG Game Benchmark at 68.3%, NetHack Depth Hits 13.2 — burny_tech · 2026-09-26
- Claude Opus proves surprisingly good at both making IK rigs and animating with them — andrew_n_carr · 2026-09-26
- Two years since o1-preview: reasoning tokens went from novelty to frontier standard — ArtificialAnlys · 2026-09-26
- Open-source JevBench adds $49/$99 paid priority benchmark runs with 48-hour turnaround — airesearch12 · 2026-09-26
- Ling Tiny 3.0 on a 2017 laptop: 8B model hits 10 tok/s, hints at edge AI future — netherreddit · 2026-09-26
- Watch Claude Opus 5.5 generate 100 matplotlib plots in one go — goodside · 2026-09-26