New paper: optimizer memory schedules outscale AdamW in overtrained Transformers

_katieeverett · x · 2026-09-10

Katie Everett and Shikai Qiu's new paper shows optimizer memory schedules (ADANA) outscale AdamW along the overtraining axis in Transformers. The thread defines the token multiplier and outscaling, then shows across 51M–253M models ADANA's advantage grows with overtraining, Muon's stays constant, and SOAP may gain at the highest OT factors.

Related event: Momentum-Scheduled ADANA Changes Scaling Exponents on the Overtraining Axis(13 posts)→

Original post →

More from Research

Research channel →