NVIDIA's MoE training recipe hits 2.21x HF BF16 baseline throughput on 8x B200 GPUs
AllThingsApx · x · 2026-09-29
An NVIDIA technical blog from the BioNeMo team details an efficient MoE training recipe for biological foundation models, benchmarking up to 2.21x the throughput of a Hugging Face BF16 baseline for Mixtral-8x7B on eight B200 GPUs. Key techniques: GroupedLinear submits multiple expert GEMMs as one grouped operation to cut kernel launch overhead; MXFP8 block-scaled 8-bit precision reduces memory vs BF16 with Blackwell hardware acceleration; and the TE Sequential API fuses GroupedLinear, ScaledSwiGLU, and routing-weight scaling into a single ForwardGroupedMLPCuTeGEMMSwiGLUMXFP8 kernel, avoiding intermediate materialization. No wet-lab data included; NVIDIA points users to the Mixtral Native Transformer Engine recipe to try it.
Related event: NVIDIA BioNeMo Boosts MoE Training Throughput 2.21x on B200 GPUs(2 posts)→
More from Infra
- Data centers are more than GPU warehouses: sovereignty means the right to switch systems off — AryHHAry · 2026-09-29
- AsideAI cuts compaction/dreaming token use 7x, doubles cache hit rate — garrytan · 2026-09-29
- Nebius cuts agent training batch collection time by 66.9%, from ~10 min to just over 3 — demian_ai · 2026-09-29
- Qwen 27B Q4 with 100K context at ~30 t/s on a 16GB AMD RX 7800 XT: full guide — Haunting-Stretch8069 · 2026-09-29
- Celesto: open-source persistent microVM computers for AI agents, boots in 500ms — aniketmaurya · 2026-09-29
- Meta Muse to cost ~$50 per user per year even under aggressive optimization, back-of-envelope says — bookwormengr · 2026-09-29