DANA/ADANA thread: momentum windows that grow with training improve scaling

_katieeverett · x · 2026-09-10

Katie Everett et al. discuss DANA/ADANA (Ferbach et al. 2025/2026), optimizers whose momentum windows grow linearly with training steps, improving scaling exponents in some regimes on power-law random features. This post asks whether a momentum schedule is even needed: longer horizons favor longer memory with smaller learning rates, and horizon-specific tuning improves AdamW baselines but doesn't fully explain ADANA's edge.

Related event: Momentum-Scheduled ADANA Changes Scaling Exponents on the Overtraining Axis(13 posts)→

Original post →

More from Research

Research channel →