CPO Optimizes Correctness-Aware Advantage Shaping in RLVR

A new study indicates that entropy is a flawed correctness signal in RLVR as it cannot distinguish between useful exploration and harmful confusion. To address this, researchers introduced Contrastive Policy Optimization (CPO) for more effective advantage shaping.

2026-07-20 ~ 2026-07-20 · 2 related posts