VC-Attention: training-free low-bit attention hits 1.9x on B200, beating FlashAttention-4
xiuyu_l · x · 2026-09-17
Nunchux AI, with researchers from MIT, CMU, UC Berkeley, Stanford and NVIDIA, released VC-Attention: fast, accurate low-bit attention without retraining.
- On MiniMax-H3 video generation (243 frames, 1344×768): 1.59x speedup on B200 and 1.51x on B300 over BF16 FlashAttention-4; the proprietary Nunchux Attention extension reaches 1.91x / 1.83x.
- Better fidelity than SageAttention2: 20.2 dB mean PSNR vs 19.9 dB across 100 prompts.
- Two innovations: V-Smooth reduces value quantization error (enabled on the first quarter of denoising steps); ExpCast-FP8 accelerates softmax.
- Compatible with existing sparse attention methods; targeted at multimodal/world-model inference where attention dominates cost.
Blog and technical report available.
More from Infra
- Nearly all top-10 PFAS makers plan production hikes to serve AI chips and data center cooling — jathansadowski · 2026-09-17
- Idle models might be smarter: a musing on GEMV vs GEMM under low concurrency — karminski3 · 2026-09-17
- University of Memphis study finds no major air quality deterioration around xAI's Colossus 1 — TinfoilTricorn · 2026-09-17
- Beam moves ~1TB every 30 minutes six months after launching on Bittensor — markjeffrey · 2026-09-17
- mlx.fast fixes speed-display bug: MLX kernels hit 80.6 tps, nearing 100% speedup milestone — HankYeomans · 2026-09-17
- IBM NorthPole claims 22x inference performance over Nvidia on 12nm process — Site-Staff · 2026-09-17