Berkeley's ComPO: zeroth-order preference alignment to mitigate likelihood displacement

Berkeley · hf · 2026-09-17

Direct preference alignment is popular for its efficiency, but likelihood displacement motivates new ways to extract signal from preference pairs with small likelihood margins. Berkeley researchers propose ComPO, a zeroth-order alignment method based on comparison oracles that extracts directional information without optimizing a differentiable preference loss on the pairs directly.

Theoretically, the paper proves convergence for the basic offline scheme under smoothness, gradient sparsity, and oracle-objective compatibility. Online ComPO keeps the offline comparison mechanism and uses unlabeled policy generations for reverse-KL control against a reference policy, with performance guarantees under local coverage and in-distribution pairwise reward accuracy.

Experiments on Mistral, Llama, Gemma-2, Qwen3, and Gemma-3 show improvements over existing direct alignment methods including length-controlled win rates, with pair-level diagnostics consistent with mitigating likelihood displacement.

Original post →

More from Research

Research channel →