MachGen Open-Sources Blackwell VC Attention: ~2x Faster Than BF16 FlashAttention on B200
MiniMax_AI · x · 2026-10-08
MachGenAI open-sourced its VC Attention variant for NVIDIA Blackwell, hitting 2x the speed of BF16 FlashAttention on B200 for models like MiniMax H3. Improvements beyond the original paper push performance 20%+ further, now within 30% of the dense attention kernels it runs in production, with far less calibration and no fine-tuning. The team frames inference optimization as kernels + recipes, arguing open kernels turn one company's moat into everyone's foundation.
More from Infra
- Dev hits 1k tokens/sec prefill at 262K context with hybrid DeepSeek V4.1 Flash build — HankYeomans · 2026-10-08
- Quantized open-weight chatbots have run fine on unaccelerated laptops for 2+ years — mattwbaker · 2026-10-08
- Kare: an on-prem GitHub Copilot gateway on Arduino VENTUNO Q with dynamic Qwen/Copilot/Foundry routing — unixterminal · 2026-10-08
- Starlink Mobile goes live in Bangladesh, connecting millions in cellular dead zones — elonmusk · 2026-10-08
- CoreWeave Launches Serverless GPUs With MicroVMs Ranging From 1 to 8 GPUs — altryne · 2026-10-08
- AMD ships ROCm 10.1, targeting the data-movement bottleneck — sn2006gy · 2026-10-08