Aleph Alpha details scaling a 30B MoE pre-training to 512 B200 GPUs at 35.3% MFU

bodonoghue85 · x · 2026-10-05

Aleph Alpha published a technical blog post, "Scaling Pre-Training in Practice: A Hierarchical Approach," detailing how they scaled pre-training of a 30B-A3B MoE model from 16 to 512 NVIDIA B200 GPUs, achieving 35.3% MFU—only 6% below perfect linear scaling.

The core method: a hierarchical approach that shrinks the hyperparameter search space at small GPU counts before scaling up, avoiding prohibitively expensive naive grid searches at 512 GPUs. The post stresses that simply buying more GPUs does not improve efficiency—untuned software stacks degrade quickly—and that frontier labs invest substantial human- and agent-hours chasing percentage-point efficiency gains.

The author also reveals that the undisclosed details connect to their open-weight model Kolibri Origin, which uses the same 30B MoE architecture. Even developers with only 1-2 GPUs can apply the optimization philosophy.

Original post →

More from Infra

Infra channel →