Beyond Entropy: CPO for correctness-aware RLVR

tencent · hf · 2026-07-20

Beyond Entropy: Correctness-Aware Advantage Shaping via Contrastive Policy Optimization

The paper argues that entropy is a poor correctness signal in reinforcement learning with verifiable rewards (RLVR) because it cannot tell useful uncertainty from harmful confusion.

It proposes Contrastive Policy Optimization (CPO), which shapes advantages using token-level contrastive disagreement between a reference-guided distribution and a vanilla generation distribution. The authors claim this disagreement is a reliable indicator of token-level correctness. They also show:

Related event: CPO Optimizes Correctness-Aware Advantage Shaping in RLVR(2 posts)→

Original post →

More from Research

Research channel →