NVIDIA’s Sol-Attn speeds up video generation inference by 2.1x

Haopeng Li · hf · 2026-07-28

NVIDIA proposes Sol-Attn for faster video generation inference

The paper targets the attention bottleneck in diffusion transformers used for high-fidelity video generation. It introduces Sol-Attn (“Sparsifying online attention”), a training-free method that combines dynamic routing, sparse computation, and approximation correction in a single online-softmax pass.

What it changes

Reported results

Related event: NVIDIA Introduces Sol-Attn to Accelerate Video Generation(2 posts)→

Original post →

More from Multimodal

Multimodal channel →