M+Adam: New Optimizer for Low-Precision Training
AnimaAnandkumar · x · 2026-07-14
The author shares the ICML paper M+Adam: Low-Precision Training via Additive–Multiplicative Optimization.
The paper proposes an optimizer combining Adam-style additive updates with Madam-style multiplicative updates to train low-precision master weights more effectively. While modern training commonly uses BF16, FP8, and FP4, it typically retains high-precision master weights; this research directly investigates optimizing on low-precision weight grids like BF16, FP8, and NVFP4.
The author's core arguments are:
- Additive updates are better suited for small weights, zeros, and sign changes.
- Multiplicative updates are more stable in large-magnitude regions and less likely to be rounded away.
- The two are complementary, effectively covering different numerical ranges.
Links to the paper and code are also included.
Related event: M+Adam Optimizer Improves Low-Precision Training(2 posts)→
More from Infra
- SkyPilot exits stealth with $20M seed round and an AI compute platform for fragmented clouds — jfiance · 2026-07-22
- Nothing phone mockup turns a film joke into a modular design meme — ZeYanjie · 2026-07-22
- Actual Computer says its inference stack is tuned for Nvidia’s consumer Blackwell lineup — markjeffrey · 2026-07-22
- Ben Bajarin says CPU demand is still being badly underestimated — BenBajarin · 2026-07-22
- An energy model says the U.S. could run short of natural gas starting in 2028 — churchkey · 2026-07-22
- Devin adds e2b sandboxes for remote agent execution — badphilosopher · 2026-07-22