Optimizers can change scaling exponents: ADANA outscales Muon, rivals SOAP on the overtraining axis

_katieeverett · x · 2026-09-10

arXiv 2609.04577 (Katie Everett et al.) studies how optimizers scale along the overtraining axis, comparing AdamW, Muon, SOAP, and the momentum-scheduled ADANA across models from 51M to 253M parameters and overtraining factors from 1x to 256x, sweeping base learning rates at every setting.

Key findings:

The author calls it the first work convincing her that optimizers can truly improve scaling exponents in Transformers, not just constant factors.

Related event: Momentum-Scheduled ADANA Changes Scaling Exponents on the Overtraining Axis(13 posts)→

Original post →

More from Models

Models channel →