Smaller Models Make Better Rejects: Study Rethinks Preference Distillation from 7B to 72B
LinkedIn · hf · 2026-10-02
LinkedIn research challenges two default assumptions in preference distillation: that self-generated failures are the most informative negatives, and that rejects must come from models at least as large as the student.
- Across students from 7B to 72B, smaller frozen models generate rejects with less inference compute yet train stronger students than self-generated rejects, on code generation and math reasoning, before and after sequence-level distillation
- The authors derive a finite-horizon utility bound for DPO in a linearized feature model that characterizes favorable reject distributions
- Three interventions: mixing rejects from smaller and student-scale models, reassigning rejects to other prompts and shuffling code tokens (still beats length-matched gibberish), and selecting lower-likelihood candidates under the reference policy
- Effective rejects preserve task structure while limiting coupling to the reference policy—and small models provide them cheaply
More from Research
- Microsoft's FOCUS compresses agent context at test time: 48% less context, +8.9 points success — dair_ai · 2026-10-02
- Audit: 23% of "wrong" cache hits in a semantic caching benchmark were identical prompts — Reasonable_Royal_621 · 2026-10-02
- Amazon Research Awards Fall 2026 opens with 9 research tracks, deadline Nov 4 — DocXavi · 2026-10-02
- Meta research: Knowledge distillation cuts training data memorization by over 50% vs fine-tuning — lasha_nlp · 2026-10-02
- Hugging Face publishes the ultimate guide to multi-harness RL — Saboo_Shubham_ · 2026-10-02
- From Transformers to MoE: The Research Papers Behind Every Major AI Concept — goyalshaliniuk · 2026-10-02