NVIDIA paper explores SOAP, Muon and other ways to scale LLM pretraining

A_K_Nain · x · 2026-07-24

SOAP, Muon, and other optimizers for scaling LLM pretraining

A new NVIDIA paper titled “SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales” focuses on how to push large-language-model pretraining to larger scales. The post points to the paper and the image shows the author list and NVIDIA affiliation.

This is research-oriented content rather than a product announcement: the key signal is the paper itself and its optimization/scaling theme.

Original post →

More from Infra

Infra channel →