ADANA paper thread part 2: caveats and confirmation of exponent-level optimizer gains
_katieeverett · x · 2026-09-10
Katie Everett wraps up part 2 of her ADANA optimizer paper thread, noting three caveats: small models (51M–253M), fixed batch of 256×2048 (ADANA may lose efficiency at larger batch), and some points requiring extrapolated baseline fits. She emphasizes this is the first time she's convinced optimizers can genuinely improve scaling exponents in Transformers, with more experiments forthcoming.
More from Research
- Bug Hunt Bench author: leaderboard noise is about 2-3 points — PawelHuryn · 2026-09-11
- Bug Hunt Bench ranks frontier coding models on 105 planted real-repo bugs — PawelHuryn · 2026-09-11
- PNAS paper shows a tiny billiard-ball system is a universal computer — undecidability lives in two dimensions — eigensteve · 2026-09-11
- New paper: Absolute pose estimation from affine cues and gravity direction — ducha_aiki · 2026-09-11
- LoMa Paper Ships REALLY HardPairs Dataset, Accepted at ECCV 2026 — ducha_aiki · 2026-09-11
- Johns Hopkins Launches Full-Stack Hands-on Robot Learning Class with SO-101 Arm Kits — _krishna_murthy · 2026-09-11