RL KL estimators may optimize the wrong direction, and k1 can collapse to zero gradient
auto_grad_ · x · 2026-07-25
- The post argues that common KL estimators used in RL losses can have incorrect gradients for the objective they are supposed to optimize.
- In particular, it claims k3 (used in DeepSeek series) is an unbiased estimator of the value of reverse KL, but its gradient is actually an unbiased estimator of the gradient of forward KL instead.
- It also points out that the original k1 estimator, the raw log-ratio, has an expectation whose gradient is zero — meaning it is effectively optimizing nothing.
- The screenshots walk through the derivations and show why the gradients line up with the opposite KL direction or collapse to zero.
Related event: New Research Reveals Flaws in KL Estimators for LLM RL Training(4 posts)→
More from Research
- Mathematician Daniel Litt Launches Problem Repo to Track Human vs AI Progress: 15 Problems, 1 Solved — littmath · 2026-09-11
- Open ECDSA.fail challenge uses AI agents to shrink Shor's-algorithm quantum circuits for Bitcoin keys — StefanoGogioso · 2026-09-11
- Alex Townsend posts 200 open problems in numerical linear algebra for humans and AI agents — IgorCarron · 2026-09-11
- Navier-Stokes, Riemann, P vs NP: what this week's math buzzwords mean for you — koltregaskes · 2026-09-11
- Fruit fly brain as an LLM: connectome-driven language model demo goes live — ngxson · 2026-09-11
- Harry Collins: LLMs can't do frontier science because they can't invent new language — whoamisri · 2026-09-11