Microsoft Research shows weak models can still distill stronger students
MicrosoftResearch · hf · 2026-08-04
Microsoft Research proposes Weak-to-Strong On-Policy Distillation (W2S-OPD), a framework that lets a strong student learn from multiple weaker models.
- It builds a proxy teacher in logit space from a positive/negative contrast pair, both smaller than the student.
- The student distills on its own rollouts with per-token reverse KL.
- The paper tests three contrast setups: post-RL vs pre-RL, larger vs smaller base model, and correct vs wrong hints.
- Across four math and three code benchmarks, W2S-OPD outperforms standard OPD, can push the student beyond the domain teacher, and still improves even when every supervision source is weaker.
- The authors also find that different contrasts emphasize different signals: some favor reasoning frameworks, others solving procedures.
More from Research
- Berkeley paper turns Gemini Robotics On-Device into a humanoid specialist via CLIFT — berkeley_ai · 2026-08-04
- LLM evaluation research says small prompt changes can flip benchmark rankings — jindong_wang92 · 2026-08-04
- A verification skill every agent needs: computer and browser use — vikvang1 · 2026-08-04
- University of Michigan lab opens five AI, ECG and multi-omics research jobs — kevinnbass · 2026-08-04
- scE2G predictions are now browsable across hundreds of cell types — anshulkundaje · 2026-08-04
- scE2G lands in Nature Genetics with a new held-out CRISPR benchmark — anshulkundaje · 2026-08-04