Aleph Alpha Details MoE Pre-training Scaling: 30B-A3B on 16-512 B200s at 35% MFU

CatAstro_Piyush · x · 2026-10-02

Aleph Alpha published a blog post detailing its hierarchical approach to pre-training scaling: the choices made at different GPU counts, taking a 30B-A3B MoE model from 16 to 512 B200 GPUs while sustaining 35% MFU and near-linear scaling. A useful engineering reference for teams running distributed MoE training.

Original post →

More from Infra

Infra channel →