New curvature-conditioned multiscale momentum optimizer significantly accelerates Muon LLM pretraining

YouJiacheng · x · 2026-09-17

An arXiv paper proposes curvature-conditioned multiscale momentum with sphere constraints for LLM pretraining. Key insight: AdamW/Muon's reliance on gradient normalization poorly mitigates ill-conditioned curvature, leaving flat eigen-directions — which dominate final loss reduction — to progress slowly.

The method applies multiscale momentum only along flat directions, pairing a slow-decay component for noise reduction with a fast-decay component for rapid curvature adaptation, plus a sphere constraint to prevent parameter inflation and overly rapid effective LR decay. Experiments show significant Muon acceleration across dense and MoE architectures at 0.12B–2.3B parameters, with theoretical verification included.

Original post →

More from Infra

Infra channel →