Aleph Alpha details scaling a 30B MoE pre-training to 512 B200 GPUs at 35.3% MFU
bodonoghue85 · x · 2026-10-05
Aleph Alpha published a technical blog post, "Scaling Pre-Training in Practice: A Hierarchical Approach," detailing how they scaled pre-training of a 30B-A3B MoE model from 16 to 512 NVIDIA B200 GPUs, achieving 35.3% MFU—only 6% below perfect linear scaling.
The core method: a hierarchical approach that shrinks the hyperparameter search space at small GPU counts before scaling up, avoiding prohibitively expensive naive grid searches at 512 GPUs. The post stresses that simply buying more GPUs does not improve efficiency—untuned software stacks degrade quickly—and that frontier labs invest substantial human- and agent-hours chasing percentage-point efficiency gains.
The author also reveals that the undisclosed details connect to their open-weight model Kolibri Origin, which uses the same 30B MoE architecture. Even developers with only 1-2 GPUs can apply the optimization philosophy.
More from Infra
- AI buildout is a doom loop: 70-80% of cloud AI revenue traces to money-losing OpenAI and Anthropic — churchkey · 2026-10-05
- Otis: an open-source AI agent that auto-sets up llama.cpp and unifies local and hosted inference — Objective-Pair8231 · 2026-10-05
- CostGraph launches GPU tracking to attribute and bill token/GPU usage for inference providers — saheedniyi_02 · 2026-10-05
- Use Magpie CLI to check quotas and route sub-agents by urgency to save tokens — lxfater · 2026-10-05
- The next AI bottleneck isn't GPUs — it's a gigawatt connected to the grid — ingliguori · 2026-10-05
- AI Is Cheap to Use, Expensive to Provide: Investor Says Compute Needs Real Spot Prices — Kyrannio · 2026-10-05