CPO: Correctness-Aware Advantage Shaping via Contrastive Policy Optimization
Weiwen Xu · hf · 2026-07-20
Background: In Reinforcement Learning with Verifiable Rewards (RLVR), entropy is commonly used for advantage shaping. However, entropy fails to distinguish useful exploration (uncertainty) from detrimental confusion, limiting its effectiveness as a correctness signal.
Method: The authors propose Contrastive Policy Optimization (CPO). It uses token-level contrastive disagreement between reference-guided and vanilla generation distributions for correctness-aware advantage shaping.
Conclusions:
- Both theoretical and empirical results show this disagreement reliably indicates token-level correctness.
- On-policy Distillation is proven to be a special case of CPO.
- CPO resolves the zero-advantage problem. Experiments demonstrate that CPO substantially outperforms entropy-based RLVR methods while maintaining strong generalization. Balancing exploitation of correct responses and exploration of incorrect ones leads to the best performance.
Related event: CPO Optimizes Correctness-Aware Advantage Shaping in RLVR(2 posts)→
More from Research
- Stanford Team Introduces Gigatoken, the World's Fastest Tokenizer — StanfordAILab · 2026-07-22
- Tabul AI launches Metal TreeSHAP to speed up Shapley values on Apple silicon — Scobleizer · 2026-07-22
- Reddit points to OpenAI’s ChatGPT Ads page — EcstaticAsparagus509 · 2026-07-22
- Open-source runtime lets each repo define its own AI code reviewer — ibabufrik · 2026-07-22
- DeepSWE: A New Benchmark for Evaluating AI Coding Agents on Real GitHub Issues — pmz · 2026-07-22
- A Rust space-economy sim runs hundreds of autonomous ships, built with Claude — kalcode · 2026-07-22