Optimizers can change scaling exponents: ADANA outscales Muon, rivals SOAP on the overtraining axis
_katieeverett · x · 2026-09-10
arXiv 2609.04577 (Katie Everett et al.) studies how optimizers scale along the overtraining axis, comparing AdamW, Muon, SOAP, and the momentum-scheduled ADANA across models from 51M to 253M parameters and overtraining factors from 1x to 256x, sweeping base learning rates at every setting.
Key findings:
- The preferred learning rate schedule can reverse across the overtraining axis; optimal weight decay scales roughly as √(OT); longer horizons favor longer fixed memory.
- ADANA's scaling advantage over AdamW persists even after tuning AdamW's fixed memory per horizon; with log-time weight decay and momentum cooldown, its exponent advantage approaches that predicted by DANA theory on power-law random features.
- Muon and SOAP provide only roughly constant token-efficiency gains over AdamW across most of the range (SOAP may extend its lead at the highest OT factors). ADANA starts behind both but closes the gap with longer training, surpassing Muon and matching SOAP at the highest OT factors.
The author calls it the first work convincing her that optimizers can truly improve scaling exponents in Transformers, not just constant factors.
More from Models
- Bug Hunt Bench ranks frontier coding models on 105 planted real-repo bugs — PawelHuryn · 2026-09-11
- GPT-6 Astra beats Factorio with enemies in 44 in-game hours at ~$4,500 API cost — liminal_bardo · 2026-09-11
- 105 hidden bugs, 2 repos: DeepSeek V4.1 Flash fixes 24 at $1.80 vs Opus 5's 27 at $51.33 — ChartsJournalX · 2026-09-11
- awesome-llm-leaderboards: an open-source directory of LLM leaderboards, pricing tables, comparison tools — Last_Establishment_1 · 2026-09-11
- Anthropic claims it works to keep eval environments unidentifiable to models — MaxKannen · 2026-09-11
- Nex N2.5 Pro released on Hugging Face with 407GB of weights — jinnyjuice · 2026-09-11