New curvature-conditioned multiscale momentum optimizer significantly accelerates Muon LLM pretraining
YouJiacheng · x · 2026-09-17
An arXiv paper proposes curvature-conditioned multiscale momentum with sphere constraints for LLM pretraining. Key insight: AdamW/Muon's reliance on gradient normalization poorly mitigates ill-conditioned curvature, leaving flat eigen-directions — which dominate final loss reduction — to progress slowly.
The method applies multiscale momentum only along flat directions, pairing a slow-decay component for noise reduction with a fast-decay component for rapid curvature adaptation, plus a sphere constraint to prevent parameter inflation and overly rapid effective LR decay. Experiments show significant Muon acceleration across dense and MoE architectures at 0.12B–2.3B parameters, with theoretical verification included.
More from Infra
- Pinterest Details the Evolution of Its Billion-Scale Embedding Retrieval Models — AxSaucedo · 2026-09-17
- OpenMuse: open-source sandboxed computer lets AI agents work without vendor lock-in — aniketmaurya · 2026-09-17
- OpenMuse: open-source sandboxed computer lets AI agents work without vendor lock-in — aniketmaurya · 2026-09-17
- Winbond to Buy Infineon's NOR Flash Business for $1.12B, Becoming World's Largest Maker — zephyr_z9 · 2026-09-17
- IonQ and ORNL demonstrate generative AI for quantum optimization — donutloop · 2026-09-17
- How do you track agent costs beyond tokens? Voice minutes and sandbox container time break cost dashboards — Any_Warning_1183 · 2026-09-17