New paper: optimizer memory schedules outscale AdamW in overtrained Transformers
_katieeverett · x · 2026-09-10
Katie Everett and Shikai Qiu's new paper shows optimizer memory schedules (ADANA) outscale AdamW along the overtraining axis in Transformers. The thread defines the token multiplier and outscaling, then shows across 51M–253M models ADANA's advantage grows with overtraining, Muon's stays constant, and SOAP may gain at the highest OT factors.
More from Research
- Bug Hunt Bench author: leaderboard noise is about 2-3 points — PawelHuryn · 2026-09-11
- Bug Hunt Bench ranks frontier coding models on 105 planted real-repo bugs — PawelHuryn · 2026-09-11
- PNAS paper shows a tiny billiard-ball system is a universal computer — undecidability lives in two dimensions — eigensteve · 2026-09-11
- New paper: Absolute pose estimation from affine cues and gravity direction — ducha_aiki · 2026-09-11
- LoMa Paper Ships REALLY HardPairs Dataset, Accepted at ECCV 2026 — ducha_aiki · 2026-09-11
- Johns Hopkins Launches Full-Stack Hands-on Robot Learning Class with SO-101 Arm Kits — _krishna_murthy · 2026-09-11