Princeton-Stanford study: optimizer rankings flip with batch size; best scaling rule is setting-dependent
xingyudang · x · 2026-10-10
A new paper and interactive blog by Xingyu Dang (Princeton), Kaiyue Wen (Stanford), and Sadhika Malladi (UCSD) systematically studies how batch size interacts with learning-rate scaling rules and optimizer choice:
- No universal scaling rule: testing hundreds of scaling rules (covering most prior recipes) with Muon, the best rule depends on the task and target batch size.
- Optimizer rankings reverse: after exhaustively retuning Adam, Lion, Muon, SOAP, and Shampoo at every batch size on modded-nanogpt track-3, MuonH beats ShampooH at 128K tokens per update, but ShampooH leads at 2M.
- Curvature-to-noise ratio (CNR): in a SignSGD toy model, higher CNR favors faster LR growth with batch size, from roughly sqrt toward linear; different directions need different scaling rules.
- SGD can beat Newton at small batches: Newton's curvature correction amplifies gradient noise, so on a noisy quadratic SGD wins small-batch and Newton wins large-batch even with optimal LR and momentum.
- Direction-aware fix: in LM pretraining, preserving small-batch updates in only the sharpest 0.001% of matrix directions removes up to 59.5% of batch-size scaling error.
An interactive blog lets readers change batch size to watch rankings flip and test their own scaling rules.
Related event: Study: Optimal Optimizer Varies With Batch Size, Rankings Can Flip(2 posts)→