NVIDIA’s Sol-Attn delivers up to 5.08× end-to-end speedup for video generation
songhan_mit · x · 2026-07-28
Sol-Attn speeds up video generation with training-free sparse attention
NVIDIA Research introduced Sol-Attn, a training-free sparse attention method that accelerates pretrained visual generators while preserving quality.
- It combines dynamic routing, sparse computation, and approximate correction in a single online-softmax pass.
- The method uses on-the-fly block thresholding for controllable budgets and reuses proxy scores to approximate skipped blocks.
- Reported speedups vs. dense FlashAttention-3:
- Wan 2.1-14B: 2.02× end-to-end
- HunyuanVideo-13B: 2.12× end-to-end
- LTX 2.3: up to 2.4× end-to-end
- Integrated into Sol-Engine with kernel fusion and caching, the gains increase further:
- Wan 2.1-14B: 3.48× end-to-end
- HunyuanVideo-13B: 5.08× end-to-end
NVIDIA says the technique is already available in Sol-Engine, and the B200 kernel is still being optimized. The paper and code are public.
More from Multimodal
- Seedance 2.0 test shows strong Genshin-style video generation with character refs — aziz4ai · 2026-07-28
- Spanish-Language AI Media Generation Video Tutorial Series Launched, from Basics to Advanced — Botoni · 2026-07-28
- Leaked Imagine Omni upgrade would bring images, video and voice into Grok Imagine — XFreeze · 2026-07-28
- ACE-Step UI claims to generate 4-minute vocal songs locally, free and open source — anselm · 2026-07-28
- A Reddit user builds a Wan 2.2 continuation workflow with LoRAs and ComfyUI — SnooMacaroons1365 · 2026-07-28
- User shares a new Midjourney style with exact prompt settings — azed_ai · 2026-07-28