NVIDIA’s SANA-Video 2.0 uses hybrid attention to generate 720p video on one GPU
nvidia · hf · 2026-07-24
SANA-Video 2.0 brings hybrid attention to video generation
NVIDIA introduces SANA-Video 2.0, a hybrid video diffusion transformer at 5B and 14B scales.
- The model is designed to generate up to 720p video on a single GPU.
- It combines gated linear attention with periodic gated-softmax anchors to recover full-rank token interactions without quadratic cost.
- Block Attention Residuals (AttnRes) reuse block summaries in later layers and raise deep-layer effective rank by about 12%.
- NVIDIA says the model is trained from scratch, not by linearizing a pretrained model.
- With 40-step sampling, SANA-Video 2.0 reaches 84.30 VBench and runs in 13.2s at 480p on one H100.
- A compiled forward pass is 3.2× faster than a matched full-softmax baseline at 720p/60s, and the full Sol-Engine stack adds another 3.58× speedup.
- The 5B pipeline is reported to run 120× faster than Wan 2.2-A14B on one H100 at the cited setting.
Overall, the paper argues that hybrid linear-softmax attention can recover much of softmax expressiveness while making long, high-resolution video generation substantially cheaper.
Related event: NVIDIA Unveils SANA-Video 2.0 for Efficient Video Generation(2 posts)→
More from Infra
- WSJ: Huawei has cut China’s foreign AI chip dependence below 60% — firstadopter · 2026-07-24
- Antirez publishes a mixed q2/q3 Laguna S2.1 quant for 64GB MacBooks — antirez · 2026-07-24
- Report says 2026 is the year edge AI shifts from pilot projects to mass deployment — 面壁智能 · 2026-07-24
- AWS outage is breaking the internet, poster says — bindureddy · 2026-07-24
- Alter Identity offers identity infrastructure for AI systems — modelcontextprotocol · 2026-07-24
- Oracle cuts 21,000 jobs as Wisconsin OpenAI campus faces a $7B power guarantee hurdle — kimmonismus · 2026-07-24