M+Adam: New Optimizer for Low-Precision Training

AnimaAnandkumar · x · 2026-07-14

The author shares the ICML paper M+Adam: Low-Precision Training via Additive–Multiplicative Optimization.

The paper proposes an optimizer combining Adam-style additive updates with Madam-style multiplicative updates to train low-precision master weights more effectively. While modern training commonly uses BF16, FP8, and FP4, it typically retains high-precision master weights; this research directly investigates optimizing on low-precision weight grids like BF16, FP8, and NVFP4.

The author's core arguments are:

Links to the paper and code are also included.

Related event: M+Adam Optimizer Improves Low-Precision Training(2 posts)→

Original post →

More from Infra

Infra channel →