Another Model Shifts to Muon Optimizer as It Emerges as the Training Default

stochasticchasm · x · 2026-09-11

A researcher observed another model adopting the Muon optimizer on a per-head basis, concluding that Muon is "becoming the default" for training.

The post links to an interesting technical breakdown of Muon, though offers few additional details. The Muon optimizer, positioned as an alternative to AdamW, appears to be moving from experiments into real production training runs.

Related event: Muon Optimizer Becomes Default as New Models Drop Adam(3 posts)→

Original post →

More from Research

Research channel →