Paper Proposes Compute-Efficient Hyperparameter Transfer for MoE

kakaocorp · hf · 2026-08-24

This paper introduces a two-step hyperparameter transfer framework designed to predict optimal learning rates for large-scale Mixture-of-Experts models. By scaling across model widths and token budgets, it enables efficient pretraining without the need for costly sweeps.

Original post →

More from Infra

Infra channel →