SANA-Video 2.0 pairs hybrid attention with 84.30 VBench and major speedups

danijarh · x · 2026-07-24

SANA-Video 2.0 is a newly released video model optimized end-to-end for efficiency while keeping quality high.

It uses a hybrid linear-softmax attention design, Block Attention Residuals, and the Sol-Engine acceleration stack. The team says it trained unified 5B and 14B models from scratch on limited resources — 16 H100 nodes for the 5B model and 48 B200 nodes for the 14B model — and reports 84.30 VBench total, 3.2× faster DiT forward passes at 720p/60s, 13.06s for 720p/5s on a single H100, and 120× faster generation than Wan 2.2-A14B under the same setup.

Original post →

More from Multimodal

Multimodal channel →