ADANA thread: momentum cooldown rule nearly doubles token multiplier vs AdamW at highest OT

_katieeverett · x · 2026-09-10

Third question in the thread: can shortening ADANA's memory near the end of training help? The authors propose a momentum cooldown rule that shrinks the window on the LR-decay timescale, nearly doubling ADANA's token multiplier relative to AdamW at the highest overtraining.

Related event: Momentum-Scheduled ADANA Changes Scaling Exponents on the Overtraining Axis(13 posts)→

Original post →

More from Research

Research channel →