KLPO: critic-free async RL for LLM agents with one rollout per prompt, no importance weights

math-ai · hf · 2026-10-08

This paper introduces KL-Regularized Policy Optimization (KLPO), addressing stale-checkpoint rollouts and trainer-sampler probability mismatch in asynchronous RL for LLM agents. Anchoring the KL regularizer at the sampler yields a closed-form Gibbs solution; KLPO fits the log-ratio optimality condition by least squares, eliminating importance weights. The authors prove unbiased gradients from Monte Carlo KL estimates, derive exact gaps of top-K/binary approximations, and show SPPO, GPO, REBEL, and BPO are special cases. Result: a critic-free update using one rollout per prompt, no group sampling needed.

Original post →

More from Research

Research channel →