ViRDM: Teacher-Free Video Distillation Hits 84.87 VBench in 16 A100-Hours
Zichong Meng · hf · 2026-09-25
ViRDM is a teacher- and critic-free post-training recipe for few-step causal video generation.
- Motivation: Existing DMD-based post-training needs a large teacher plus an online critic to estimate distributional discrepancies
- Method: Transfers representation distribution matching (RDM) from one-step image generation to few-step causal video, solving memory-intractable gradient paths, a distinct video optimization regime, and underconstrained temporal dynamics via stochastically truncated clean-exit supervision, a lightweight VAE decoder, staged vector–Jacobian products, and lightweight dynamics regularization
- Results: Turns three-network distillation into generator-only training; with just 20 generator updates and 16 A100 GPU-hours it reaches 84.87 on VBench, beating the prior best few-step causal baseline by 0.36, with exploratory results for 1/2/4-step bidirectional generation
More from Multimodal
- HuggingChat's ML Intern trains a LoRA from one prompt in ~75 minutes — Gradio · 2026-09-25
- Gradio demos LoRA novel-view editing: 25s renders, evals on 160 held-out objects — Gradio · 2026-09-25
- AI video as a game engine: PixVerse R2 streams interactive worlds from one causal backbone — future_coded · 2026-09-25
- Adding "micro expressions" to prompts yields subtler AI character emotions — R34vspec · 2026-09-25
- User generates a music video from old material with Opus 5.5 — repligate · 2026-09-25
- Same prompt, Opus 5.5 one-shot video generation put to a public retest with different tools — drrickio · 2026-09-25