ADANA outscales AdamW across overtraining in 51M–253M Transformers

_katieeverett · x · 2026-09-10

Core result of the thread: across 51M–253M parameter Transformers, ADANA outscales AdamW as the overtraining factor grows. Matrix-preconditioned optimizers lead at small OT; Muon's advantage is roughly constant across OT, while SOAP may gain further at the highest OT factors tested.

Related event: Momentum-Scheduled ADANA Changes Scaling Exponents on the Overtraining Axis(13 posts)→

Original post →

More from Research

Research channel →