SANA-Video 2.0 pairs hybrid attention with 84.30 VBench and major speedups
danijarh · x · 2026-07-24
SANA-Video 2.0 is a newly released video model optimized end-to-end for efficiency while keeping quality high.
It uses a hybrid linear-softmax attention design, Block Attention Residuals, and the Sol-Engine acceleration stack. The team says it trained unified 5B and 14B models from scratch on limited resources — 16 H100 nodes for the 5B model and 48 B200 nodes for the 14B model — and reports 84.30 VBench total, 3.2× faster DiT forward passes at 720p/60s, 13.06s for 720p/5s on a single H100, and 120× faster generation than Wan 2.2-A14B under the same setup.
More from Multimodal
- Midjourney prompt mixes 1970s retro art with infographic-style mechanical diagrams — michaelrabone · 2026-07-24
- A full GPT image 2 prompt template turns ChatGPT into a watercolor portrait generator — SimplyAnnisa · 2026-07-24
- The real shift in generative AI is conversational workflows, not one-shot prompts — LawfulnessNext3503 · 2026-07-24
- FLUX 3 Video enters Early Access, but production users still lack pricing and reliability data — MembershipEmergency7 · 2026-07-24
- Nova Crew sci-fi video features a cinematic spaceship flying past a planet — HomeRemedyHealer · 2026-07-24
- ListenHub and Labnana add Seedream 5.0 Pro with image editing — oran_ge · 2026-07-24