NVIDIA Paper Shows SOAP and Muon Optimizers Outperform AdamW in LLM Pretraining
An NVIDIA paper reveals that advanced optimizers like SOAP and Muon outperform the traditional AdamW in large-scale LLM pretraining. The research highlights that AdamW hits a wall with large batch sizes, making higher-order optimizers crucial for scaling.
2026-07-24 ~ 2026-07-24 · 3 related posts
- NVIDIA paper explores SOAP, Muon and other ways to scale LLM pretraining — A_K_Nain · 2026-07-24
- NVIDIA study says SOAP and Muon beat AdamW on multi-billion-parameter LLM pretraining — sudoraohacker · 2026-07-24
- Notes on SOAP, Muon and KL-SOAP: AdamW hits a large-batch wall at 72B scale — tokenbender · 2026-07-24