RL KL estimators may optimize the wrong direction, and k1 can collapse to zero gradient
auto_grad_ · x · 2026-07-25
- The post argues that common KL estimators used in RL losses can have incorrect gradients for the objective they are supposed to optimize.
- In particular, it claims k3 (used in DeepSeek series) is an unbiased estimator of the value of reverse KL, but its gradient is actually an unbiased estimator of the gradient of forward KL instead.
- It also points out that the original k1 estimator, the raw log-ratio, has an expectation whose gradient is zero — meaning it is effectively optimizing nothing.
- The screenshots walk through the derivations and show why the gradients line up with the opposite KL direction or collapse to zero.
Related event: New Research Reveals Flaws in KL Estimators for LLM RL Training(4 posts)→
More from Research
- AI slop is already clogging PR review and weakening the credit system behind science — rbhar90 · 2026-07-27
- ICML 2026 oral paper replication scores stay middling after a stricter re-scoring — profjamesevans · 2026-07-27
- Long-running agents will need immutable event logs, this thread argues — sebpaquet · 2026-07-27
- Seed IQ navigates Doom II, prompting questions about benchmarks beyond ARC-AGI — Fit_Transition8824 · 2026-07-27
- Agentic Data Science in Practice: Agents Write Code but Answer Wrong Questions — hugobowne · 2026-07-27
- A concise canon of foundational papers in ML, systems, NLP, speech, and audio — deliprao · 2026-07-27