VC-Attention: Training-Free Low-Bit Attention Speeds Up Inference on B200
Researchers from MIT, CMU, Berkeley, Stanford and NVIDIA introduced VC-Attention, a training-free low-bit attention method that tackles Value outlier and softmax bottlenecks, delivering 1.6x attention speedup on B200 (surpassing FlashAttention-4) and up to 3.6x faster video generation.
2026-09-17 ~ 2026-09-17 · 3 related posts
- VC-Attention: training-free low-bit attention hits 1.9x on B200, beating FlashAttention-4 — xiuyu_l · 2026-09-17
- VC-Attention: Training-Free Low-Bit Attention Speeds Video Diffusion Up to 3.6x — nunchux-inc · 2026-09-17
1 near-duplicate retellings: songhan_mit