RIPO Overcomes PPO-Clip's Exploration Collapse in LLM RL

A new paper reveals that PPO-Clip causes exploration collapse in LLM reinforcement learning due to a flawed Euclidean metric. The proposed RIPO algorithm fixes this Riemannian manifold issue, boosting AIME24 performance by 60%.

2026-07-23 ~ 2026-07-25 · 2 related posts