Ai2 Releases Olmo-core 3: Open MoE Training Framework Scales to 128 Experts with <5% Throughput Loss
allen_ai · x · 2026-10-01
Ai2 released Olmo-core 3, a redesigned open mixture-of-experts training system built to scale MoE training toward the trillion-parameter range. It underpins the next generation of Olmo.
- Benchmark: growing the expert pool from 8 to 128 (still 4 experts per token, 3.2B active params) raised total capacity from 4.6B to 47B with less than 5% throughput loss; the infrastructure has been benchmarked beyond 1T total parameters
- Motivation: expert routing and coordination costs erode MoE efficiency advantages at scale
- The tech report shares experiments, findings, and failures, including "token gerrymandering," where a load-balancing score could improve as workloads became less balanced
Related event: Ai2 Open-Sources Olmo-core 3 for Trillion-Parameter MoE Training(2 posts)→
More from Infra
- AMD shows off Helios system with OpenAI aboard amid deepening infra ties — AnushElangovan · 2026-10-01
- Qwen-Image 2.1 prompt enhancer hits 4.4x speedup in ComfyUI, now runs on 8GB VRAM — mozophe · 2026-10-01
- Running Omarchy desktop in Windows via WSL with GPU acceleration and 4K multi-monitor support — sytelus · 2026-10-01
- Trader initiates Cerebras position, betting SRAM-based inference beats HBM as agents multiply model calls — Sethwinterroth · 2026-10-01
- Cerebras bull case: OpenAI paid tier, ~750 tok/s, $20B+ potential value and $25B RPO — Sethwinterroth · 2026-10-01
- MLX-Serve 26.10.1 ships with up to 66% faster Qwen3.8 27B inference on Apple Silicon — TheMoonMidas · 2026-10-01