ByteDance's DMAD hits FID 1.04 one-step on ImageNet, SOTA few-step visual generation
ByteDance · hf · 2026-10-08
ByteDance released DMAD (Distribution Matching as Adversarial Distillation), a new method for fast visual generation that removes DMD's costly auxiliary diffusion model.
Core idea: Instead of fitting a separate score model to the student's evolving distribution, DMAD recasts distribution matching as classification — two discriminator heads on a shared backbone separate real data, teacher samples, and student samples, and linear losses on their logits directly learn the log-density ratios. The paper proves these losses recover the DMD distribution-matching gradient at the discriminator optimum, plus a gap-based reweighting scheme that adapts teacher supervision across noise levels.
Results:
- FID 1.04 with one-step generation on ImageNet-64x64
- FID 14.47 with four-step SDXL on COCO-10K
- VBench total 85.15 with four-step Wan2.1-T2V-14B, beating compared few-step methods and multi-step teachers
- On MiniMax-H3-33B joint audio-video generation, 79.1% human preference over DMD2 and 84.6% over rCM (excluding ties)
Code, models, and demos are open-sourced.
More from Multimodal
- Concept art to rigged character: skill partitioning + ComfyUI + Trellis pipeline demoed — majidmanzarpour · 2026-10-09
- AA-Video-T2V v2.0 leaderboard: Wan 3.0 tops, Grok Imagine Video 1.5 Lite debuts — ArtificialAnlys · 2026-10-09
- Grok Imagine 1.5 Lite by Use Case: Best at Architecture & Real Estate, Worst at Film — ArtificialAnlys · 2026-10-09
- Grok Imagine Video 1.5 Lite Beats Veo 3.1 on AA Leaderboard at a Third the Price — ArtificialAnlys · 2026-10-09
- Krea 2 Prompt Writer tutorial walks through the workflow — Intelligent-While146 · 2026-10-09
- Open ComfyUI Node Pack: Krea 2, Z-Image, MiniMax Music 3 and Apple MLX LLMs — TimeTruth2490 · 2026-10-09