One prompt is enough: distillation from a single query hits 71.5% of full-data gains
burkov · x · 2026-09-06
Researchers from Tsinghua, UCAS, Northeastern, UIUC and Johns Hopkins show on-policy distillation — where a teacher corrects the student at every token — recovers most of the improvement from 17,000 training queries using just one query, reaching 71.5% of the state regions visited by full-data training. A single prompt spawns many prompt-response "states," each giving the teacher another correction opportunity; 16 semantically diverse queries extend coverage further.
More from Research
- ICML paper: hyperfitting a LoRA on final 5 layers removes AI slop, code released — grimjim · 2026-09-06
- $1,000-trained HRM-Text shows Sapient bet on recurrence before OpenAI's Astra — rohanpaul_ai · 2026-09-06
- Entropy trajectory shape predicts Qwen3-4B errors and transfers to unseen tasks — Happy_Brilliant7827 · 2026-09-06
- Pedro Domingos quips: just wait until you see AI churn out string theory papers — pmddomingos · 2026-09-06
- Lean formalization of agent foundations papers catches an error in Logical Induction — jessi_cata · 2026-09-06
- 17-page PDF offers a clear introduction to Hidden Markov Models — Roger_M_Taylor · 2026-09-06