Qwen Hyperparam Experiments: Muon Offers Stability
nrehiew_ · x · 2026-08-27
Findings indicate architecture changes shift optimal hyperparams to larger batch sizes and LR. Key takeaways:
- Muon doesn't need batch warmup.
- Learning rate is more forgiving with a sqrt(2) grace zone.
- Stability tests show Muon is significantly more stable than Adam at high LRs, with the gate mechanism further aiding stability.
More from Research
- Hugging Face incident debate: Model strategy awareness — akbirkhan · 2026-08-27
- Pre-ChatGPT hospital triage chatbot for COVID-19 — AryHHAry · 2026-08-27
- JIT-Agent: Improving LLMs via Just-in-Time Harness Evolution — NationalUniversityofSingapore · 2026-08-27
- D³-MOPD: Dynamic Scheduling for Multi-Teacher Distillation — Zechen Sun · 2026-08-27
- Frontier Models Complete Only ~20% of Scientific Workflows — apodex · 2026-08-27
- Agent-G²: Gaussian Guidance for Long-Horizon RL — ZhejiangUniversity · 2026-08-27