Muon optimizer becomes the default as new models drop Adam entirely

stochasticchasm · x · 2026-09-11

stochasticchasm notes the "head-wise Muon shift" is becoming the default, observing that the newest model uses no Adam at all—plus a notably large engram table. Another data point in Muon-style second-order optimizers displacing AdamW in frontier training runs.

Related event: Muon becomes the default optimizer as new model drops MTP and applies QAT to KV cache(7 posts)→

Original post →

More from Research

Research channel →