MachGen Open-Sources Blackwell VC Attention: ~2x Faster Than BF16 FlashAttention on B200

MiniMax_AI · x · 2026-10-08

MachGenAI open-sourced its VC Attention variant for NVIDIA Blackwell, hitting 2x the speed of BF16 FlashAttention on B200 for models like MiniMax H3. Improvements beyond the original paper push performance 20%+ further, now within 30% of the dense attention kernels it runs in production, with far less calibration and no fine-tuning. The team frames inference optimization as kernels + recipes, arguing open kernels turn one company's moat into everyone's foundation.

Original post →

More from Infra

Infra channel →