Muon optimizer becomes the default as new models drop Adam entirely
stochasticchasm · x · 2026-09-11
stochasticchasm notes the "head-wise Muon shift" is becoming the default, observing that the newest model uses no Adam at all—plus a notably large engram table. Another data point in Muon-style second-order optimizers displacing AdamW in frontier training runs.
More from Research
- k3 Report Section Confirms Millions of Concurrent Sandboxes in Its RL Training Run — stochasticchasm · 2026-09-11
- Persimmon unveiled: first large-scale model to simulate human conversation — niloofar_mire · 2026-09-11
- ApprenticeBench: Top Models Now Beat APIs Through GUIs, the 'CUA Tax' Has Disappeared — ysu_nlp · 2026-09-11
- k3 RL Run Reportedly Used ~50M Sandboxes With Millions Concurrent, Checkpoints Merged Across Scaffolds — stochasticchasm · 2026-09-11
- World models vs LLMs: why next-token prediction still lacks internal representations — TheTuringPost · 2026-09-11
- ECDSA.fail challenge paper on arXiv: AI agents optimize quantum circuits for Bitcoin's secp256k1 — jedisct1 · 2026-09-11