Marin's 535B MoE hero run hits ~27% MFU on 11 NVL72 racks: expert parallelism deep dive

dlwh · x · 2026-10-03

First in a series on how Marin's 535B-parameter MoE hero run reached 27% MFU across 11 NVL72 racks. Expert parallelism let them grow the model 50% from the planned 360B to 535B total (23B active per token). Their fixed-capacity EP kernel achieved 21.5% MFU with a 3.4% token drop rate. The post candidly covers what they tried (FSDP, off-the-shelf and custom kernels), what failed, the ugly solution they launched with, and the better one swapped in mid-run.

Original post →

More from Infra

Infra channel →