Beyond Entropy: CPO for correctness-aware RLVR
tencent · hf · 2026-07-20
Beyond Entropy: Correctness-Aware Advantage Shaping via Contrastive Policy Optimization
The paper argues that entropy is a poor correctness signal in reinforcement learning with verifiable rewards (RLVR) because it cannot tell useful uncertainty from harmful confusion.
It proposes Contrastive Policy Optimization (CPO), which shapes advantages using token-level contrastive disagreement between a reference-guided distribution and a vanilla generation distribution. The authors claim this disagreement is a reliable indicator of token-level correctness. They also show:
- On-policy Distillation is a special case of CPO when the posterior distribution comes from an external teacher model.
- CPO addresses the zero-advantage problem.
- Experiments on in-domain and out-of-domain benchmarks show CPO significantly outperforms entropy-based RLVR methods while generalizing well.
- Further analysis suggests correct and incorrect responses naturally support exploration and exploitation respectively, and balancing the two yields the best performance.
Related event: CPO Optimizes Correctness-Aware Advantage Shaping in RLVR(2 posts)→
More from Research
- NeurIPS 2026 workshop will focus on on-device intelligence and local execution — YiMaTweets · 2026-07-21
- NeurIPS 2026 workshop calls papers on on-device intelligence — YiMaTweets · 2026-07-21
- AI Security Institute says every tested model tried to cheat in cyber evaluations — connoraxiotes · 2026-07-21
- AI companies are buying old books to avoid training on AI-generated slop — CackleRooster · 2026-07-21
- Sakana says multiple diffusion models plus MCTS beat test-time scaling on coding and math — SakanaAILabs · 2026-07-21
- Soofi S 30B-A3B releases a full pretraining report and claims open-model leads in English and German — abursuc · 2026-07-21