NVIDIA’s SANA-Video 2.0 uses hybrid attention to generate 720p video on one GPU
nvidia · hf · 2026-07-24
SANA-Video 2.0 brings hybrid attention to video generation
NVIDIA introduces SANA-Video 2.0, a hybrid video diffusion transformer at 5B and 14B scales.
- The model is designed to generate up to 720p video on a single GPU.
- It combines gated linear attention with periodic gated-softmax anchors to recover full-rank token interactions without quadratic cost.
- Block Attention Residuals (AttnRes) reuse block summaries in later layers and raise deep-layer effective rank by about 12%.
- NVIDIA says the model is trained from scratch, not by linearizing a pretrained model.
- With 40-step sampling, SANA-Video 2.0 reaches 84.30 VBench and runs in 13.2s at 480p on one H100.
- A compiled forward pass is 3.2× faster than a matched full-softmax baseline at 720p/60s, and the full Sol-Engine stack adds another 3.58× speedup.
- The 5B pipeline is reported to run 120× faster than Wan 2.2-A14B on one H100 at the cited setting.
Overall, the paper argues that hybrid linear-softmax attention can recover much of softmax expressiveness while making long, high-resolution video generation substantially cheaper.
Related event: NVIDIA Unveils SANA-Video 2.0 for Efficient Video Generation(2 posts)→
More from Infra
- Nvidia Is Now Core to Every Major Robotaxi Stack at Commercial Scale — pdamodaran · 2026-09-11
- 12 KV Cache Reduction Techniques Every AI Engineer Should Understand, Explained — blaizedsouza · 2026-09-11
- The shadow GPU capacity market is formalizing, with Meta selling excess compute to outside buyers — DavidLinthicum · 2026-09-11
- Engram's random reads don't suit SSDs; CPU-memory over NVLink could serve all 72 GPUs — bookwormengr · 2026-09-11
- 80% of the DIY LLM inference hype posters have already quit — it's brutally hard systems work — abhijithneil · 2026-09-11
- Hugging Face's Ultra Scale Playbook: a free book on training LLMs on GPU clusters — mdancho84 · 2026-09-11