Muon Optimizer Beats Adam on Small Superhuman Agents
A Kimi researcher found Muon significantly outperforms Adam when training small superhuman agents on games, drones, driving and business simulations. A related hypothesis suggests exploration-trained RL agents benefit more from Muon than LLM pretraining does.
2026-08-29 ~ 2026-08-29 · 2 related posts
- Muon massively beats Adam on tiny superhuman agents across games, drones, driving — jsuarez · 2026-08-29
- Hypothesis: Strong RL Agents via Free Exploration Produce More Diverse Behaviors — YouJiacheng · 2026-08-29