Beyond Entropy: CPO for correctness-aware RLVR
tencent · hf · 2026-07-20
Beyond Entropy: Correctness-Aware Advantage Shaping via Contrastive Policy Optimization
The paper argues that entropy is a poor correctness signal in reinforcement learning with verifiable rewards (RLVR) because it cannot tell useful uncertainty from harmful confusion.
It proposes Contrastive Policy Optimization (CPO), which shapes advantages using token-level contrastive disagreement between a reference-guided distribution and a vanilla generation distribution. The authors claim this disagreement is a reliable indicator of token-level correctness. They also show:
- On-policy Distillation is a special case of CPO when the posterior distribution comes from an external teacher model.
- CPO addresses the zero-advantage problem.
- Experiments on in-domain and out-of-domain benchmarks show CPO significantly outperforms entropy-based RLVR methods while generalizing well.
- Further analysis suggests correct and incorrect responses naturally support exploration and exploitation respectively, and balancing the two yields the best performance.
Related event: CPO Optimizes Correctness-Aware Advantage Shaping in RLVR(2 posts)→
More from Research
- Nature paper images cellular activity across all organs, revealing body-wide circuits — arjunrajlab · 2026-09-11
- SignNet 1M Dataset Released for Sign Language Research — ducha_aiki · 2026-09-11
- ECCV26 Oral: Flow Matching Enables Single-Stage Multi-View Point Cloud Registration — ducha_aiki · 2026-09-11
- InFlux++ Method Released — ducha_aiki · 2026-09-11
- Skyfall GS Uses Flux to Refine Gaussian Splatting, Accepted at ECCV 2026 — ducha_aiki · 2026-09-11
- Could 10k agents discover learning methods beyond backprop, or just tweak existing ones? — SeunghyunSEO7 · 2026-09-11