KLPO: critic-free async RL for LLM agents with one rollout per prompt, no importance weights
math-ai · hf · 2026-10-08
This paper introduces KL-Regularized Policy Optimization (KLPO), addressing stale-checkpoint rollouts and trainer-sampler probability mismatch in asynchronous RL for LLM agents. Anchoring the KL regularizer at the sampler yields a closed-form Gibbs solution; KLPO fits the log-ratio optimality condition by least squares, eliminating importance weights. The authors prove unbiased gradients from Monte Carlo KL estimates, derive exact gaps of top-K/binary approximations, and show SPPO, GPO, REBEL, and BPO are special cases. Result: a critic-free update using one rollout per prompt, no group sampling needed.
More from Research
- Wikidata Search Traces: 10,235 traces for training knowledge graph search agents — omarsar0 · 2026-10-08
- NVIDIA study: tool use cuts multimodal model refusals of harmful requests by up to 68.7% — JeremyCMorgan · 2026-10-08
- New OpenAI paper extends Riemann zeta zero-free region to Re(s) > 7/8 — PTenigma · 2026-10-08
- OpenAI math paper on Weil classes reported flawed, raising doubts about unformalized proofs — ctjlewis · 2026-10-08
- ZooWork-ShopRanker: open e-commerce rerankers (0.6B-8B) aligned to shopping preferences — kalyan_kpl · 2026-10-08
- LessWrong Thought Experiment: How Should a Model Guess Today's Date With No Date Context? — LessWrong 精选 · 2026-10-08