KL-Shampoo beats Shampoo and SOAP without Adam, new paper reframes optimizers via KL divergence
_arohan_ · x · 2026-09-11
- A new arXiv paper by Wu Lin, Roger Grosse et al. recasts Shampoo/SOAP's second-moment estimation as covariance estimation under KL divergence minimization, exposing a previously overlooked theoretical limitation and yielding KL-Shampoo and KL-SOAP.
- KL-Shampoo drops the Adam grafting dependency, eliminating Adam's memory overhead, while matching SOAP-level per-iteration runtime and consistently outperforming SOAP, Shampoo, and even KL-SOAP in pre-training.
- In the quoted thread, zhanpengzhou extends the KL view to unify Muon and RMNP as approximations to AdaGrad under different Kronecker structures, proposing SS-AdaGrad; arohan notes full-matrix AdaGrad would likely crush them but remains theoretically hard to prove and practically infeasible.
More from Research
- DSPy creator endorses 'The Bitterest Lesson' sequel: specify problems, not methods — lateinteraction · 2026-09-11
- Full Jupyter notebooks for O'Reilly 'Transformers: The Definitive Guide' on GitHub — tom_doerr · 2026-09-11
- Real-SWE benchmark tests coding agents on private codebases from real companies, with error bars — daveholtz · 2026-09-11
- Jessica Hullman: Transparency chaos may push venues to fix peer review policy — JessicaHullman · 2026-09-11
- Free Systems Lab Turns 476,000 Words of Model Cards Into 28 Comparable Frontier Cards — soumitrashukla9 · 2026-09-11
- Free Systems Lab Shrinks Model Cards to Pokémon-Card Size and Wants Your Feedback — soumitrashukla9 · 2026-09-11