Meta's Muse Realtime Avatar research blog: audio-driven DiT brings any image to life in real time

AIatMeta · x · 2026-09-24

Meta published its research blog on Muse Realtime Avatar. Key points: (1) any reference media — portraits, illustrations, animals, objects — becomes an expressive live conversational avatar with coherent facial, hand and full-body motion; (2) Muse Realtime Voice produces a speech-token (VQ) stream carrying content and prosody, which Avatar consumes to generate synchronized video; (3) it's an audio-driven Diffusion Transformer generating causal video chunks conditioned on tokens, reference media and a rolling window of recent latents, keeping computation bounded indefinitely; (4) companion tweets cover a 2-step distilled student and blind-test wins over commercial systems.

Related event: Meta Unveils Muse Realtime Avatar: Sub-second Realtime Digital Humans(18 posts)→

Original post →

More from Multimodal

Multimodal channel →