NVIDIA BioNeMo recipe boosts MoE biological model training up to 2.21x on eight B200s
AllThingsApx · x · 2026-10-01
NVIDIA's technical blog details an efficient MoE training recipe for biological foundation models using BioNeMo and Transformer Engine:
- GroupedLinear: submits multiple expert GEMMs as one grouped operation, removing kernel-launch overhead from per-expert Python loops.
- MXFP8: block-scaled 8-bit precision cuts memory vs BF16 with Blackwell hardware acceleration.
- Kernel fusion: the Sequential API fuses GroupedLinear, ScaledSwiGLU, and routing-weight scaling into a single kernel, avoiding intermediate materialization.
- Benchmark: up to 2.21x throughput over the Hugging Face baseline on eight B200 GPUs. A Mixtral Native Transformer Engine recipe is available to try.
Related event: NVIDIA BioNeMo Speeds Up MoE Training Up to 2.21x Over HF Baseline(3 posts)→
More from Infra
- Cloudflare launches agent-themed batch: pay-per-use gateway, AutoRouter, 6x faster containers — threepointone · 2026-10-01
- Distributed compute market adds 10 B300 nodes, rentable from a single node — markjeffrey · 2026-10-01
- 27B at Q5 with full 131k context on one 24GB RTX 3090, 13-17% faster — bjivanovich · 2026-10-01
- netkit paper: container network namespaces cost up to 31% throughput on Linux — tianyin_xu · 2026-10-01
- Dev sorts PCIe issues, spins up 7-GPU local rig with 6000 Pro and 6000 ADA cards — TheZachMueller · 2026-10-01
- Ben Bajarin: HPE and others use 'scale-out' vs 'scale-up' in conflicting ways — BenBajarin · 2026-10-01