New Research Reveals Flaws in KL Estimators for LLM RL Training

Recent preprints reveal that common reverse KL estimators used in LLM reinforcement learning, such as the K3 estimator used in DeepSeek, provide incorrect gradients. The studies show that K1 can suffer from near-zero gradients, highlighting critical implementation pitfalls in RL training.

2026-07-25 ~ 2026-07-26 · 4 related posts