CPO: Correctness-Aware Advantage Shaping via Contrastive Policy Optimization

Weiwen Xu · hf · 2026-07-20

Background: In Reinforcement Learning with Verifiable Rewards (RLVR), entropy is commonly used for advantage shaping. However, entropy fails to distinguish useful exploration (uncertainty) from detrimental confusion, limiting its effectiveness as a correctness signal.

Method: The authors propose Contrastive Policy Optimization (CPO). It uses token-level contrastive disagreement between reference-guided and vanilla generation distributions for correctness-aware advantage shaping.

Conclusions:

Related event: CPO Optimizes Correctness-Aware Advantage Shaping in RLVR(2 posts)→

Original post →

More from Research

Research channel →