Muon optimizer eliminates need for batch size warmup

nrehiew_ · x · 2026-08-27

Loss curves reveal that using Muon alone results in very high gradient norms and residual activations. Findings indicate that architecture and optimizer changes shift optimal hyperparameters towards significantly larger batch sizes and learning rates. Key takeaways: batch size warmup is not needed with Muon due to large penalties at small sizes, and learning rates are more forgiving with a grace zone.

Related event: Qwen Training Details: Full-Network Muon Optimizer and Residual Stream Design(5 posts)→

Original post →

More from Research

Research channel →