Muon nearly doubles agentic RL success in GiGPO, but only with the right setup
rohanpaul_ai · x · 2026-07-22
Muon can substantially improve agentic reinforcement learning, but only under the right training setup.
- The paper compares vanilla Muon against AdamW on sparse-reward agentic RL using matched single-seed comparisons.
- With GiGPO, which estimates advantage from repeated states, Muon raises late-window validation success from 0.290 to 0.546 — an 88%+ relative gain.
- The authors find the effect depends on both the advantage estimator and learning rate: at 3e-5, Muon improves GRPO from 0.161 to 0.268; at 1e-5, the gap narrows as GraphGPO’s late-window performance nearly saturates.
- They conclude that an optimizer should not be judged in isolation; the surrounding RL setup can determine whether it helps or hurts.
More from Research
- New health-time-series benchmark shows LLMs lag behind classic ML baselines — yang_yuzhe · 2026-07-22
- AutoLab benchmark shows frontier models win long-horizon tasks by persisting, not guessing — rohanpaul_ai · 2026-07-22
- Report says Gemini Flash keeps parity while cutting per-task cost by 30–40% — sujingshen · 2026-07-22
- Delineate Anything v2 maps field boundaries across 61 countries with a 73M-instance dataset — Mykola Lavreniuk · 2026-07-22
- KernelBench adds a CUDA-only sub-benchmark and reruns Fable on RTX PRO 6000, H100, and B200 — TheZachMueller · 2026-07-22
- Niantic Spatial, Flexion and NVIDIA push humanoid sim-to-real work — ExtensionEcho3 · 2026-07-22