NVIDIA's SANA-Video 2.0: Hybrid Attention Model Runs 720p on a Single RTX 5090
mmowg · reddit · 2026-08-03
NVIDIA has released SANA-Video 2.0, a new Video Diffusion Transformer available in 5B and 14B parameter versions. It features a redesigned architecture with a 3:1 Hybrid Linear-Softmax Attention, combining the speed of linear attention with the expressiveness of softmax.
Powered by Sol-Engine optimization (kernel fusion, sparse attention, TensorRT, MXFP4/MXFP8), it achieves a 3.58x speedup. It is the first NVIDIA video model designed for consumer GPUs, capable of generating 720p video on a single RTX 5090. It scores 84.30 on VBench and is up to 120x faster than Wan 2.2-A14B on the same hardware.
Currently, the model is "research-open" with no code, weights, or licensing terms published yet.
More from Multimodal
- Same Prompt Image Generation Comparison: REVE 2.1 vs Meta AI — LudovicCreator · 2026-08-03
- Help: How Should Text Tags Be Written for Training WAN 2.2 Video LoRA? — Overall-Reporter-440 · 2026-08-03
- Creating a Fantasy Movie Trailer Entirely with Hailuo AI — menhguin · 2026-08-03
- MiniMax H3 Open Weights Released: Community Tests Keyframe Interpolation — Hannibalj2ca · 2026-08-03
- Adding a Hint of Gravity Lets the AI Model Draw a Heart — johnowhitaker · 2026-08-03
- ComfyUI Core Dev Dismisses Rumors: Hunyuan3D 2.0 Open-Source Release Still On — OneTrueTreasure · 2026-08-03