VC-Attention: Training-Free Low-Bit Attention Speeds Up Inference on B200

Researchers from MIT, CMU, Berkeley, Stanford and NVIDIA introduced VC-Attention, a training-free low-bit attention method that tackles Value outlier and softmax bottlenecks, delivering 1.6x attention speedup on B200 (surpassing FlashAttention-4) and up to 3.6x faster video generation.

2026-09-17 ~ 2026-09-17 · 3 related posts

1 near-duplicate retellings: songhan_mit