OKLS brings KL-optimal Shampoo to language model training with 1.45× parameter efficiency

aryaman2020 · x · 2026-07-29

OKLS proposes a KL-optimal Shampoo variant for language model training

The paper introduces Online KL Shampoo (OKLS), an optimizer that approximates full-matrix AdaGrad more closely than diagonal methods, while staying practical for large-scale language model training.

Original post →

More from Research

Research channel →