Momentum-Scheduled ADANA Changes Scaling Exponents on the Overtraining Axis
Katie Everett and Shikai Qiu released a new paper (arXiv 2609.04577) systematically studying optimizer scaling laws along the overtraining axis. The core finding: an optimizer's momentum memory schedule doesn't just improve constant factors—it can genuinely change the scaling exponent, i.e., it "outscales" AdamW.
Confirmed
- The authors build on the ADANA family of methods (from Ferbach et al. 2025/2026's DANA/ADANA line): the momentum window grows linearly with training steps, improving performance in certain regimes under a power-law random features setup
- In Transformer experiments along the overtraining axis, momentum scheduling consistently beats AdamW, beats Muon, and approaches SOAP
- Setup: models of 51M–253M parameters, fixed batch (256×2048), comparing AdamW, Muon, SOAP, and ADANA
- The authors clarified background in the thread, including their definition of overtraining (a token multiplier)
Not confirmed (author-stated limitations)
- Model scale is small (51M–253M); conclusions at larger scale remain to be verified
- Batch is fixed at 256×2048; ADANA may lose its efficiency advantage earlier as batch grows
- Some comparison points require extrapolating the baselines
Why it matters
- Overtraining (smaller models consuming more tokens) is now standard practice; if optimizer choice can change the scaling exponent on this axis rather than just the constant factor, that means real-world training gains
- This is evidence pushing optimizer research from "a bit faster at the same exponent" to "a different exponent altogether," though scale and batch conditions remain limited
2026-09-10 ~ 2026-09-10 · 13 related posts
Primary sources
- [source] New paper: optimizer memory schedules can outscale AdamW's exponent on the overtraining axis — _katieeverett · 2026-09-10
- New paper: optimizer memory schedules outscale AdamW in overtrained Transformers — _katieeverett · 2026-09-10
- [source] ADANA outscales AdamW across overtraining in 51M–253M Transformers — _katieeverett · 2026-09-10
- Token multiplier: 2x means the baseline needs twice the tokens — _katieeverett · 2026-09-10
- Defining outscaling: token multiplier rising along the overtraining axis — _katieeverett · 2026-09-10
- DANA/ADANA thread: momentum windows that grow with training improve scaling — _katieeverett · 2026-09-10
- ADANA thread: horizon-specific momentum tuning can't explain ADANA's edge over AdamW — _katieeverett · 2026-09-10
- ADANA thread: log-time weight decay helps AdamW and SOAP but hurts Muon at high overtraining — _katieeverett · 2026-09-10
- ADANA thread: momentum cooldown rule nearly doubles token multiplier vs AdamW at highest OT — _katieeverett · 2026-09-10
- ADANA thread: DANA theory says log-time memory outscales fixed memory with slope 2-kappa — _katieeverett · 2026-09-10
- ADANA thread: cooldown plus log-time weight decay puts AdamW's equivalent OT on a 1.15 slope — _katieeverett · 2026-09-10
- [source] ADANA paper thread part 2: caveats and confirmation of exponent-level optimizer gains — _katieeverett · 2026-09-10
- Optimizers can change scaling exponents: ADANA outscales Muon, rivals SOAP on the overtraining axis — _katieeverett · 2026-09-10