MoE Training Cost Cut: Proxy Models Predict Optimal Learning Rates at Trillion-Token Scale
burny_tech · x · 2026-08-25
Addressing the prohibitive cost of hyperparameter tuning for trillion-token Mixture-of-Experts (MoE) models, this paper presents a compute-efficient, two-step transfer framework.
- Methodology: The approach first uses Maximal Update Parameterization (μP) to transfer optimal learning rates across model widths. It then establishes a predictive scaling law to extrapolate these rates from short training runs to massive token horizons (e.g., 10T tokens).
- Validation: The method achieved high fidelity (R²=0.95) in extrapolating ideal learning rates for large-scale training using data from small proxy models. The authors successfully pretrained a 155B MoE model (17B active params) using configurations predicted from much cheaper runs.
- Impact: This validates that optimal configurations for full-scale target models can be accurately predicted with minimal ablation costs, eliminating the need for expensive sweeps at scale.
More from Research
- Thinking Machines proposes a safe path for open-weight model releases — luke_drago_ · 2026-08-25
- Headlong experiments with persistent agency via exponential backoff — lateinteraction · 2026-08-25
- CoRL 2026 Workshop on Memory for Robot Foundation Models CFP — EricLengyel · 2026-08-25
- Researchers shifting from alternative architectures to inference optimization — eigenron · 2026-08-25
- Jcode bench introduces first uncontaminatable open benchmark — ycombinator · 2026-08-25
- Ai2 and UW Seek Participants for Study on Overseeing Long-Horizon Agents — ChengleiSi · 2026-08-25