RIPO says PPO-Clip’s Euclidean metric is collapsing exploration in LLM RL

GenSI · hf · 2026-07-23

RIPO argues PPO-Clip is geometrically mismatched for LLM RL

This paper claims that PPO-Clip fails in LLM reinforcement learning because it measures policy updates with a Euclidean metric, which does not match the intrinsic geometry of the policy Riemannian manifold. The authors argue that this mismatch makes updates too conservative in low-probability regions and too aggressive in high-probability regions, leading to exploration collapse.

Proposed method

Reported results

The core claim is that the issue is not just a heuristic tuning problem, but a deeper geometric mismatch in the clipping objective.

Related event: RIPO Overcomes PPO-Clip's Exploration Collapse in LLM RL(2 posts)→

Original post →

More from Research

Research channel →