New research: standard SGD matches AdamW for LLM RL training, with far less memory overhead

zhaoran_wang · x · 2026-09-14

A new study challenges the default assumption that AdamW is required for training LLMs with reinforcement learning.

Key findings:

Original post →

More from Infra

Infra channel →