ADANA paper thread part 2: caveats and confirmation of exponent-level optimizer gains

_katieeverett · x · 2026-09-10

Katie Everett wraps up part 2 of her ADANA optimizer paper thread, noting three caveats: small models (51M–253M), fixed batch of 256×2048 (ADANA may lose efficiency at larger batch), and some points requiring extrapolated baseline fits. She emphasizes this is the first time she's convinced optimizers can genuinely improve scaling exponents in Transformers, with more experiments forthcoming.

Related event: Momentum-Scheduled ADANA Changes Scaling Exponents on the Overtraining Axis(13 posts)→

Original post →

More from Research

Research channel →