Cursor Open-Sources MoE Training Megakernel, Claiming ~40% E2E Speedup on B200
Dany0 · reddit · 2026-08-06
Cursor Open-Sources MoE Training Megakernel
The Cursor team has open-sourced a faster megakernel for MoE (Mixture of Experts) training, specifically optimized for B200 GPUs.
Performance
- Official Claims: Reports an end-to-end (e2e) speedup of 40% and a 140% speedup for forward passes.
- Author's Take: The poster cautions against blindly trusting benchmarks and suggests running personal tests. They estimate that compared to a naive kernel anyone could write, the actual e2e speedup is likely in the 10-20% range.
The project is released under the Apache 2.0 open-source license and is free for the community.
Related event: Cursor Open-Sources MoK MoE Megakernel, Nearly Doubling Compute(7 posts)→
More from Infra
- Chamath Warns: AI Token Bill Doubles Every 45 Days While Productivity Grows Just 5% — rohanpaul_ai · 2026-08-06
- SanDisk projects NAND market revenue to exceed $300 billion in 2026 — Beth_Kindig · 2026-08-06
- Local AI Hardware Guide: Choosing Between RTX 50-series and AMD for MiniMax H3 — Eden1506 · 2026-08-06
- Samsung to Lock 60-70% of Production in Long-Term Deals, Tech Giants as Key Clients — Beth_Kindig · 2026-08-06
- SanDisk Executives Assert: Over 80% Gross Margin is a 'Fair Return' — firstadopter · 2026-08-06
- Google's AI token processing surges 330x in two years, signaling booming inference demand — Beth_Kindig · 2026-08-06