Berkeley's ComPO: zeroth-order preference alignment to mitigate likelihood displacement
Berkeley · hf · 2026-09-17
Direct preference alignment is popular for its efficiency, but likelihood displacement motivates new ways to extract signal from preference pairs with small likelihood margins. Berkeley researchers propose ComPO, a zeroth-order alignment method based on comparison oracles that extracts directional information without optimizing a differentiable preference loss on the pairs directly.
Theoretically, the paper proves convergence for the basic offline scheme under smoothness, gradient sparsity, and oracle-objective compatibility. Online ComPO keeps the offline comparison mechanism and uses unlabeled policy generations for reverse-KL control against a reference policy, with performance guarantees under local coverage and in-distribution pairwise reward accuracy.
Experiments on Mistral, Llama, Gemma-2, Qwen3, and Gemma-3 show improvements over existing direct alignment methods including length-controlled win rates, with pair-level diagnostics consistent with mitigating likelihood displacement.
More from Research
- XConf Estimates LLM Confidence From the Model's Own Track Record, Beating Self-Consistency at 1/10 Cost — CambUni · 2026-09-17
- Researchers Uncover 'Value Flattening' in PPO Critics — Sparse Supervision on 3 States Fixes It — Shanghai-AI-Laboratory · 2026-09-17
- PANORAMA Grounds Every Caption Phrase to Pixel-Level Masks, Tops New PanoCaps Benchmark — Panorama-grounding · 2026-09-17
- Four Scheduling Techniques Flatten MoE Training Memory Peaks, Enabling 1M Context at 10.4x Throughput — Shrey Pandit · 2026-09-17
- Extending SGD diffusion approximations to optimization over probability distributions — burkov · 2026-09-17
- Mallat's one-slide test: are generative models generalizing or just memorizing? — prof_kamilov · 2026-09-17