NVIDIA’s PDD speeds image and video generation with 4–8-step distillation
nvidia · hf · 2026-07-29
- NVIDIA’s paper proposes Parallel Decoding Distillation (PDD), a trajectory-based distillation method to speed up diffusion and flow-matching models for image and video generation.
- The method is designed to be simpler and more scalable than VSD/adversarial-loss-based approaches, which are hard to optimize and can collapse diversity.
- PDD predicts multiple denoising steps per network evaluation, supports a variable number of function evaluations (NFE), and works with any pre-trained model.
- The authors report SOTA results at 4–8 NFE on LTX-2.3 Text-to-Video/Audio, Wan 14B Text-to-Video, and Qwen-Image Text-to-Image, with better video diversity as well.
More from Multimodal
- Flux 3 prompt aims for a bootleg phone recording of a metal encore — fofrAI · 2026-07-29
- Interactive keyboard-controlled video model runs in real time with just 0.5B parameters — multimodalart · 2026-07-29
- VisoMaster open-source tool swaps faces in images and videos with LivePortrait tuning — tom_doerr · 2026-07-29
- A generated img 2.0 meme shows Elon Musk talking with Sam Altman — SkyNo7576 · 2026-07-29
- ComfyUI 0.29 adds streaming video, GPT 5.6, Claude Opus 5, and new partner nodes — pytraveler · 2026-07-29
- Creator says an album was made end to end with AI using Suno — Kyrannio · 2026-07-29