Open-source marin project pushes single-rack training MFU from 28.2% to 31.1%
dlwh · x · 2026-10-04
A PR in the marin-community repo raises single-rack (GB200 NVL72, 64 GPUs) MoE training MFU from 28.24% to 31.13%, cutting the moeheroep step from 13.90s to 12.61s with loss within run-to-run noise. A final commit adds a PGLE profile builder built from the program's own trace, reaching 12.52s (31.37% MFU). Gains came from two campaigns: working down the most expensive scopes in the step, and benchmarking OLMo-core 3's BF16 MoE kernels against this path and keeping the wins. Benchmarks restore from step 180000, fixed seed, median of 100 steps.
More from Infra
- Investors want to fund "insurance for AI" startups wrapping data center leases — katieruthmishra · 2026-10-04
- Data tower waste heat could warm 100,000 homes: 134 MW at 85°C, ~1 TWh/year — IgorCarron · 2026-10-04
- Ex-OpenAI researcher flags: SPV shills from six months ago are now selling you data centers — suchenzang · 2026-10-04
- MTPLX Ships New Version: Cache Copies Cut to One, Much Faster Decoding Past 140k Tokens — HankYeomans · 2026-10-04
- Samsung: HBM to consume 30% of DRAM wafer capacity by 2027, up from ~20% — Beth_Kindig · 2026-10-04
- Nebius up 160% vs CoreWeave's 10%: the deciding factor isn't revenue or backlog — Beth_Kindig · 2026-10-04