Marin's 535B MoE hero run hits ~27% MFU on 11 NVL72 racks: expert parallelism deep dive
dlwh · x · 2026-10-03
First in a series on how Marin's 535B-parameter MoE hero run reached 27% MFU across 11 NVL72 racks. Expert parallelism let them grow the model 50% from the planned 360B to 535B total (23B active per token). Their fixed-capacity EP kernel achieved 21.5% MFU with a 3.4% token drop rate. The post candidly covers what they tried (FSDP, off-the-shelf and custom kernels), what failed, the ugly solution they launched with, and the better one swapped in mid-run.
More from Infra
- mamf-finder adds FP8/MXFP4/NVFP4 support for real GPU TFLOPS benchmarking — StasBekman · 2026-10-03
- Measured on B200: nvfp4 is ~9% more efficient than mxfp4 with higher accuracy — pick nvfp4 on Blackwell — StasBekman · 2026-10-03
- LithosAI launches LithosBox: millisecond snapshot-and-fork sandboxes for AI agents — JiaZhihao · 2026-10-03
- RTX Spark laptops and mini desktops rumored Oct 7 launch, $1800-$2900 with 24GB-128GB — Porespellar · 2026-10-03
- Wish list: a Qwen4 27B with 100B+ Engram offloaded to RAM and NVMe for local users — casper_hansen_ · 2026-10-03
- State of Local AI 2026: one gaming GPU now matches the world's best model from Feb 2026 — Scobleizer · 2026-10-03