Zero-Data DPO: Crafting Rejected Responses via Random Prompts
natashajaques · x · 2026-07-07
Introduces a simple method that delivers significant performance gains across a range of benchmarks on fine-tuned models, requiring no extra data, labels, or supervision.
The approach constructs DPO preference pairs: the preferred response yc is generated using the real prompt x, while the rejected response yr is generated using a random prompt x'. The author will reveal more details on Wednesday.
More from Research
- Researcher bootstraps from fly connectome to build increasingly intelligent connectomes — airkatakana · 2026-09-11
- CellFluxRL: RL-based biological grounding for virtual cell models, submitted to ECCV 2026 — Prof_Lundberg · 2026-09-11
- PiPNN nearest-neighbor search wins three awards, up to 78x faster index building — khademinori · 2026-09-11
- Steerable Visual Representations Presented as ICML Long Oral — y_m_asano · 2026-09-11
- OpenCVL: a satellite-to-photo registration dataset at ECCV 2026 — ducha_aiki · 2026-09-11
- Diverse VPR work submitted to ECCV 2026 — ducha_aiki · 2026-09-11