NVIDIA's MoE training recipe hits 2.21x HF BF16 baseline throughput on 8x B200 GPUs

AllThingsApx · x · 2026-09-29

An NVIDIA technical blog from the BioNeMo team details an efficient MoE training recipe for biological foundation models, benchmarking up to 2.21x the throughput of a Hugging Face BF16 baseline for Mixtral-8x7B on eight B200 GPUs. Key techniques: GroupedLinear submits multiple expert GEMMs as one grouped operation to cut kernel launch overhead; MXFP8 block-scaled 8-bit precision reduces memory vs BF16 with Blackwell hardware acceleration; and the TE Sequential API fuses GroupedLinear, ScaledSwiGLU, and routing-weight scaling into a single ForwardGroupedMLPCuTeGEMMSwiGLUMXFP8 kernel, avoiding intermediate materialization. No wet-lab data included; NVIDIA points users to the Mixtral Native Transformer Engine recipe to try it.

Related event: NVIDIA BioNeMo Boosts MoE Training Throughput 2.21x on B200 GPUs(2 posts)→

Original post →

More from Infra

Infra channel →