SparkDiffusion from PKU/Tsinghua/Alibaba hits 265x video generation speedup on a single RTX 5090
机器之心 · wechat · 2026-09-28
Researchers from Peking University, Tsinghua, and Alibaba released SparkDiffusion — the first unified open-source framework combining sparse attention, few-step distillation, and FP8 quantization to accelerate DiT video generation, with weights and full training code released.
- Key finding, the "High-Sparsity Trap": at 97% sparsity, training loss keeps dropping while video quality collapses (broken figures, temporal chaos), and longer training can't fix it. Oracle intervention experiments show the root cause: structural errors in high-noise steps get amplified along the denoising trajectory.
- Solution: Terminal-Aligned Supervision constrains what each step ultimately generates; RoLA sparse attention (block-sparse plus low-rank linear branch restoring global context) and CrossDistill trajectory-mixed distillation (PCM at high noise for diversity, DMD at low noise to fix endpoint errors) yield a 3-step CFG-free student, plus FP8 quantization.
- Results: 265x speedup on a single RTX 5090 (720P-14B, 5s video in 18s); 220x on H100 with latency down to 8s; VBench-2.0 only drops from 60.2 to 59.77.
- Why it matters: gains scale with resolution (O(L²) attention), bringing high-quality video generation to consumer GPUs; supports Wan2.1/2.2 T2V and I2V, with extensions to autoregressive and omni-modal generation planned.
More from Multimodal
- Open-source Valen: a visual decision model scoring candidates in 122–128 ms — aigclink · 2026-09-28
- NotebookLM can now generate faceless videos in seconds: 7 prompts to do it all — anthara_ai · 2026-09-28
- One Prompt, Full Lesson: Opus 5.5 Removes the Production Bottleneck in Educational Video — victor_explore · 2026-09-28
- 63 Claude Opus motion graphics videos hit 13M views in a week, all one prompt away — tibo_maker · 2026-09-28
- GPT-6 Astra demo turns real-room video into interactive 3D worlds for robot training — 141_1337 · 2026-09-28
- Quantum Neon Origami: a glowing geometric-fold prompt template for image generation — LudovicCreator · 2026-09-28