One prompt is enough: distillation from a single query hits 71.5% of full-data gains

burkov · x · 2026-09-06

Researchers from Tsinghua, UCAS, Northeastern, UIUC and Johns Hopkins show on-policy distillation — where a teacher corrects the student at every token — recovers most of the improvement from 17,000 training queries using just one query, reaching 71.5% of the state regions visited by full-data training. A single prompt spawns many prompt-response "states," each giving the teacher another correction opportunity; 16 semantically diverse queries extend coverage further.

Original post →

More from Research

Research channel →