New paper: optimizer memory schedules can outscale AdamW's exponent on the overtraining axis

_katieeverett · x · 2026-09-10

A new paper by Katie Everett and Shikai Qiu shows optimizer memory schedules can outscale AdamW across the overtraining (OT) axis in Transformers — the first convincing evidence that optimizers can improve scaling exponents, not just constants.

Related event: Momentum-Scheduled ADANA Changes Scaling Exponents on the Overtraining Axis(13 posts)→

Original post →

More from Research

Research channel →