Muon optimizer eliminates need for batch size warmup
nrehiew_ · x · 2026-08-27
Loss curves reveal that using Muon alone results in very high gradient norms and residual activations. Findings indicate that architecture and optimizer changes shift optimal hyperparameters towards significantly larger batch sizes and learning rates. Key takeaways: batch size warmup is not needed with Muon due to large penalties at small sizes, and learning rates are more forgiving with a grace zone.
More from Research
- Hugging Face incident debate: Model strategy awareness — akbirkhan · 2026-08-27
- Pre-ChatGPT hospital triage chatbot for COVID-19 — AryHHAry · 2026-08-27
- JIT-Agent: Improving LLMs via Just-in-Time Harness Evolution — NationalUniversityofSingapore · 2026-08-27
- D³-MOPD: Dynamic Scheduling for Multi-Teacher Distillation — Zechen Sun · 2026-08-27
- Frontier Models Complete Only ~20% of Scientific Workflows — apodex · 2026-08-27
- Agent-G²: Gaussian Guidance for Long-Horizon RL — ZhejiangUniversity · 2026-08-27