VC-Attention: 1.6x attention speedup on B200 without retraining, from MIT/Stanford/NVIDIA team
songhan_mit · x · 2026-09-17
Researchers from MIT, CMU, UC Berkeley, Stanford and NVIDIA introduce VC-Attention, targeting the softmax bottleneck that limits low-bit attention for fast video generation on B200/B300 GPUs:
- ExpCast-FP8: a simple linear mapping to FP8 codes that bypasses the expensive exponential and cast operations in softmax.
- V-Smooth: reduces value quantization error, achieving better fidelity than SageAttention2.
On MiniMax-H3 it delivers 1.6x attention speedup on B200 and 1.5x on B300 over BF16 FlashAttention-4; the proprietary Nunchux Attention extension reaches 1.9x/1.8x. No retraining needed and it composes with existing sparse attention methods, moving multimodal inference toward faster, cheaper serving. Blog and technical report are available.
Related event: VC-Attention: Training-Free Low-Bit Attention Speeds Up Inference on B200(3 posts)→
More from Infra
- Jensen Huang says safety-testing data centers will become AI's third compute demand pillar — rohanpaul_ai · 2026-09-17
- Jensen Huang: a 1GW NVIDIA AI factory costs $50-60B but generates ~$50B in annual rental revenue — rohanpaul_ai · 2026-09-17
- Ilya warns neoclouds' weak cybersecurity invites rogue AI agents to hijack compute — Miles_Brundage · 2026-09-17
- AMD's free AI Developer Program: $100 cloud credits, Discord access, hardware raffles — wkmyrhang · 2026-09-17
- After AWS me-central-1 loss, dev jokes about explaining the outage to Codex weekly — andersonbcdefg · 2026-09-17
- The RAM Crisis Is Only Just the Beginning as AI Demand Squeezes Supply — perelin · 2026-09-17