Aleph Alpha Details MoE Pre-training Scaling: 30B-A3B on 16-512 B200s at 35% MFU
CatAstro_Piyush · x · 2026-10-02
Aleph Alpha published a blog post detailing its hierarchical approach to pre-training scaling: the choices made at different GPU counts, taking a 30B-A3B MoE model from 16 to 512 B200 GPUs while sustaining 35% MFU and near-linear scaling. A useful engineering reference for teams running distributed MoE training.
More from Infra
- AI Capex Now Exceeds the Railroad Boom's GDP Share, Yet Demand Lags — FinanceYF5 · 2026-10-02
- Real-time AI serving costs up to 56x more than needed — one engine hits 56 sessions per H100 — Ok_boss_labrunz · 2026-10-02
- Cloudflare launches SQL API to query Workers logs and traces, replacing GraphQL plans — irvinebroque · 2026-10-02
- Cloudflare launches Web Search API via AI Gateway with Exa, Linkup and Ceramic — michellechen · 2026-10-02
- Lightpanda 1.0 ships: a Zig-built browser for AI agents, out of beta — jedisct1 · 2026-10-02
- Qwen3.8-27B coder quant fits a 24GB GPU with 262k context at 40 t/s — W61k3r · 2026-10-02