Zero-Data DPO: Crafting Rejected Responses via Random Prompts

natashajaques · x · 2026-07-07

Introduces a simple method that delivers significant performance gains across a range of benchmarks on fine-tuned models, requiring no extra data, labels, or supervision.

The approach constructs DPO preference pairs: the preferred response yc is generated using the real prompt x, while the rejected response yr is generated using a random prompt x'. The author will reveal more details on Wednesday.

Original post →

More from Research

Research channel →