Paper finds common KL estimators in LLM RL can produce the wrong gradients
burny_tech · x · 2026-07-26
- A paper on KL regularization in RL training of LLMs argues that common sample-based KL estimators can produce the wrong gradients.
- The authors show that the k3 estimator used in DeepSeek-style setups is unbiased for the value of expected reverse KL, but its gradient is not the gradient of KL; instead, it behaves like an unbiased estimator for the gradient of expected forward KL.
- Their empirical study finds that in on-policy settings, biased-gradient estimator choices can destabilize training, while unbiased-gradient configurations improve both in-domain and out-of-domain performance.
- They also note that KL regularization can help stabilize asynchronous training and provide practical takeaways for RL post-training of LLMs.
Related event: New Research Reveals Flaws in KL Estimators for LLM RL Training(4 posts)→
More from Research
- 3D ResNet Paper Crosses 3,000 Citations Eight Years After CVPR 2018 — HirokatuKataoka · 2026-09-11
- Jeff Heaton's Intro to the Math of Neural Networks eBook Is Free to Download — blaizedsouza · 2026-09-11
- Mathematician Daniel Litt Launches Problem Repo to Track Human vs AI Progress: 15 Problems, 1 Solved — littmath · 2026-09-11
- Open ECDSA.fail challenge uses AI agents to shrink Shor's-algorithm quantum circuits for Bitcoin keys — StefanoGogioso · 2026-09-11
- Alex Townsend posts 200 open problems in numerical linear algebra for humans and AI agents — IgorCarron · 2026-09-11
- Navier-Stokes, Riemann, P vs NP: what this week's math buzzwords mean for you — koltregaskes · 2026-09-11