VC-Attention: 1.6x attention speedup on B200 without retraining, from MIT/Stanford/NVIDIA team

songhan_mit · x · 2026-09-17

Researchers from MIT, CMU, UC Berkeley, Stanford and NVIDIA introduce VC-Attention, targeting the softmax bottleneck that limits low-bit attention for fast video generation on B200/B300 GPUs:

On MiniMax-H3 it delivers 1.6x attention speedup on B200 and 1.5x on B300 over BF16 FlashAttention-4; the proprietary Nunchux Attention extension reaches 1.9x/1.8x. No retraining needed and it composes with existing sparse attention methods, moving multimodal inference toward faster, cheaper serving. Blog and technical report are available.

Related event: VC-Attention: Training-Free Low-Bit Attention Speeds Up Inference on B200(3 posts)→

Original post →

More from Infra

Infra channel →