CPO: Correctness-Aware Advantage Shaping via Contrastive Policy Optimization
Weiwen Xu · hf · 2026-07-20
Background: In Reinforcement Learning with Verifiable Rewards (RLVR), entropy is commonly used for advantage shaping. However, entropy fails to distinguish useful exploration (uncertainty) from detrimental confusion, limiting its effectiveness as a correctness signal.
Method: The authors propose Contrastive Policy Optimization (CPO). It uses token-level contrastive disagreement between reference-guided and vanilla generation distributions for correctness-aware advantage shaping.
Conclusions:
- Both theoretical and empirical results show this disagreement reliably indicates token-level correctness.
- On-policy Distillation is proven to be a special case of CPO.
- CPO resolves the zero-advantage problem. Experiments demonstrate that CPO substantially outperforms entropy-based RLVR methods while maintaining strong generalization. Balancing exploitation of correct responses and exploration of incorrect ones leads to the best performance.
Related event: CPO Optimizes Correctness-Aware Advantage Shaping in RLVR(2 posts)→
More from Research
- 3D ResNet Paper Crosses 3,000 Citations Eight Years After CVPR 2018 — HirokatuKataoka · 2026-09-11
- Sample selection and ordering matter a lot in LLM training: DataFlex makes data scheduling dynamic — Puzzleheaded_Box2842 · 2026-09-11
- Jeff Heaton's Intro to the Math of Neural Networks eBook Is Free to Download — blaizedsouza · 2026-09-11
- Mathematician Daniel Litt Launches Problem Repo to Track Human vs AI Progress: 15 Problems, 1 Solved — littmath · 2026-09-11
- Open ECDSA.fail challenge uses AI agents to shrink Shor's-algorithm quantum circuits for Bitcoin keys — StefanoGogioso · 2026-09-11
- Alex Townsend posts 200 open problems in numerical linear algebra for humans and AI agents — IgorCarron · 2026-09-11