Muon Optimizer Beats Adam on Small Superhuman Agents

A Kimi researcher found Muon significantly outperforms Adam when training small superhuman agents on games, drones, driving and business simulations. A related hypothesis suggests exploration-trained RL agents benefit more from Muon than LLM pretraining does.

2026-08-29 ~ 2026-08-29 · 2 related posts