Meta's Muse Realtime Avatar research blog: audio-driven DiT brings any image to life in real time
AIatMeta · x · 2026-09-24
Meta published its research blog on Muse Realtime Avatar. Key points: (1) any reference media — portraits, illustrations, animals, objects — becomes an expressive live conversational avatar with coherent facial, hand and full-body motion; (2) Muse Realtime Voice produces a speech-token (VQ) stream carrying content and prosody, which Avatar consumes to generate synchronized video; (3) it's an audio-driven Diffusion Transformer generating causal video chunks conditioned on tokens, reference media and a rolling window of recent latents, keeping computation bounded indefinitely; (4) companion tweets cover a 2-step distilled student and blind-test wins over commercial systems.
Related event: Meta Unveils Muse Realtime Avatar: Sub-second Realtime Digital Humans(18 posts)→
More from Multimodal
- Claude Opus 5.5 generates a full launch video — animation, music, voiceover — in 20 minutes — cedric_chee · 2026-09-26
- Feeding YuE 2 an empty lyric field yields a song full of gibberish vocals — SteveLittleFish · 2026-09-26
- "Opus 5.5" Rumored Release Draws Rave First Impressions and Cynicism — chaumian · 2026-09-26
- Reddit user shares striking clip: 'Video models are getting good' — we_are_mammals · 2026-09-26
- Gemma plays Snake straight from pixels via VLM gateway, under 240ms p99 at ~$0.00007/image — spillai · 2026-09-26
- Opus 5.5 directs a sci-fi short on the Arecibo message via Krea MCP and Hyperframes — angrypenguinPNG · 2026-09-26