CPO Optimizes Correctness-Aware Advantage Shaping in RLVR
A new study indicates that entropy is a flawed correctness signal in RLVR as it cannot distinguish between useful exploration and harmful confusion. To address this, researchers introduced Contrastive Policy Optimization (CPO) for more effective advantage shaping.
2026-07-20 ~ 2026-07-20 · 2 related posts
- CPO: Correctness-Aware Advantage Shaping via Contrastive Policy Optimization — Weiwen Xu · 2026-07-20
- Beyond Entropy: CPO for correctness-aware RLVR — tencent · 2026-07-20