How Muse streams live video: 40-step diffusion teacher distilled into 2-step causal student

AIatMeta · x · 2026-09-24

Meta details Muse's live video streaming approach: a 40-step diffusion teacher with 3-way CFG (120 evals per video chunk) is distilled into an unguided 2-step causal student with a fixed-length KV cache, achieving near-teacher quality with 60x fewer evaluations; self-forcing helps the student resist drift over long conversations.

Related event: Meta Unveils Muse Realtime Avatar: Sub-second Realtime Digital Humans(18 posts)→

Original post →

More from Multimodal

Multimodal channel →