Hyperball optimizer wrapper boosts Muon pretraining speed by 20-30%
teortaxesTex · x · 2026-08-11
New research proposes Hyperball, an optimizer wrapper designed to address the issue of diminishing gains for matrix-based optimizers like Muon as model scale increases.
- Mechanism: Hyperball sets the Frobenius norms of weight matrices and their corresponding optimizer updates to fixed constants.
- Results: On Qwen3-style models up to 1.2B parameters, Muon Hyperball achieves a 20-30% token equivalent speedup over weight decay baselines.
- Benefits: The method also improves learning rate transfer across different widths and depths.
- Theory: The approach is motivated by theory showing that training with weight decay leads to an equilibrium weight norm depending solely on training hyperparameters, which then dictates the angular learning rate.
Related event: Hyperball Optimizer Breaks Muon Scaling Bottleneck for 30% Speedup(2 posts)→
More from Research
- Maglev: Sliding Recurrent Memory for Parallel Training and Sequential Decoding — ChengleiSi · 2026-08-12
- NeurIPS 2026 to Host Inaugural Workshops Dedicated to Diffusion Language Models — CSProfKGD · 2026-08-12
- How to Get Training Data as a Hobbyist: Extracting Ebooks from LibGen — theshawwn · 2026-08-12
- Princeton & NVIDIA Joint Paper Proposes L0-L5 Autonomy Levels for Future AI Labs — MengdiWang10 · 2026-08-12
- New Paper Shows Multi-stage SFT Causes Catastrophic Forgetting While RL Excels — joecole · 2026-08-12
- Codex Falsely Reports Success 4.1% of the Time, Open-Source Tool Reveals — Due_Emu_8229 · 2026-08-12