New paper: optimizer memory schedules can outscale AdamW's exponent on the overtraining axis
_katieeverett · x · 2026-09-10
A new paper by Katie Everett and Shikai Qiu shows optimizer memory schedules can outscale AdamW across the overtraining (OT) axis in Transformers — the first convincing evidence that optimizers can improve scaling exponents, not just constants.
- Horizon-specific momentum tuning improves baselines but doesn't explain ADANA's edge over AdamW.
- Log-time weight decay helps AdamW, ADANA, and SOAP but harms Muon at high OT; best treatment depends on optimizer and horizon.
- A momentum cooldown rule (shortening memory near end of training on the LR-decay timescale) nearly doubles ADANA's token multiplier vs AdamW at the highest OT.
- DANA theory predicts log-time memory outscales fixed memory with slope 2-κ (κ≈0.85 for language data); combined, AdamW's equivalent OT roughly follows a 1.15 slope.
- Caveats: small models (51M–253M), fixed batch of 256×2048 sequences, some points extrapolated.
More from Research
- Steerable Visual Representations Presented as ICML Long Oral — y_m_asano · 2026-09-11
- OpenCVL: a satellite-to-photo registration dataset at ECCV 2026 — ducha_aiki · 2026-09-11
- Diverse VPR work submitted to ECCV 2026 — ducha_aiki · 2026-09-11
- EASE: evidence-anchored spatial attention lifts multimodal RLVR by up to 3.1 points, EMNLP 2026 — jiqizhixin · 2026-09-11
- P=NP Explained: Why Class Schedules and Circuit Routing Are the Real Hard Problems — thesaraharminta · 2026-09-11
- Hypothesis: ASI Has a Mathematical Incentive to Preserve Human Diversity — No_Cause_2731 · 2026-09-11