VC-Attention: Training-Free Low-Bit Attention Speeds Video Diffusion Up to 3.6x

nunchux-inc · hf · 2026-09-17

For diffusion transformers, long spatiotemporal sequences make attention the dominant deployment cost, and low-bit quantization hits two walls: value outliers with no fixed channel/spatiotemporal structure dominate output error, and the high-precision softmax exponential becomes the longest pipeline stage on datacenter GPUs.

VC-Attention is a training-free framework pairing:

Implemented for B200, B300, H200, RTX PRO 6000, and RTX 5090. Across Wan2.2, LongCat-Video, HunyuanVideo-1.5, and MiniMax-H3, it beats low-bit baselines on fidelity and speeds the attention kernel 1.46–1.59x over BF16 FlashAttention-4 on datacenter Blackwell/Hopper and 2.3–3.6x on workstation cards; end-to-end clip generation is 1.13–1.19x and 1.36–1.70x faster respectively.

Related event: VC-Attention: Training-Free Low-Bit Attention Speeds Up Inference on B200(3 posts)→

Original post →

More from Infra

Infra channel →