LLM RL Not Generalizing? Researcher Points to Wrong KL Optimization

cjmaddison · x · 2026-07-30

In response to the view that reinforcement learning (RL) for LLMs isn't generalizing as expected by labs, a researcher pointed out that this phenomenon might be due to optimizing KL divergence the wrong way round.

Original post →

More from Research

Research channel →