How Muse streams live video: 40-step diffusion teacher distilled into 2-step causal student
AIatMeta · x · 2026-09-24
Meta details Muse's live video streaming approach: a 40-step diffusion teacher with 3-way CFG (120 evals per video chunk) is distilled into an unguided 2-step causal student with a fixed-length KV cache, achieving near-teacher quality with 60x fewer evaluations; self-forcing helps the student resist drift over long conversations.
Related event: Meta Unveils Muse Realtime Avatar: Sub-second Realtime Digital Humans(18 posts)→
More from Multimodal
- Claude Opus 5.5 generates a full launch video — animation, music, voiceover — in 20 minutes — cedric_chee · 2026-09-26
- Feeding YuE 2 an empty lyric field yields a song full of gibberish vocals — SteveLittleFish · 2026-09-26
- "Opus 5.5" Rumored Release Draws Rave First Impressions and Cynicism — chaumian · 2026-09-26
- Reddit user shares striking clip: 'Video models are getting good' — we_are_mammals · 2026-09-26
- Gemma plays Snake straight from pixels via VLM gateway, under 240ms p99 at ~$0.00007/image — spillai · 2026-09-26
- Opus 5.5 directs a sci-fi short on the Arecibo message via Krea MCP and Hyperframes — angrypenguinPNG · 2026-09-26