AI Researchers Debate Whether Distillation Drives China's Model Gains
On September 25, AI researchers xeophon and fleetwood clashed over multiple rounds on X about the role of distillation in Chinese models closing the gap with US models, becoming one of the day's hot topics in AI circles.
Confirmed
- xeophon stated flatly that distillation is irrelevant to the big picture and far from the reason Chinese models are approaching US ones, and he wished this argument would die. He pushed back on the claim that "strong models' gains mainly come from distillation," calling it outdated: SFT as the main source of capability gains was a 2023-24 story, while today's progress mainly comes from RL environments.
- On the objection that the distillation path "hits a floor" — i.e., some model must actually complete the task, or you must buy trajectories from human annotators — xeophon countered that during training you can first use hints to guide the model through the initial prompt, then remove those hints from the SFT data; besides, there are plenty of open-source models for generating trajectory data.
- fleetwood took a different stance: he sees distillation as a necessary but not sufficient condition for training strong models, with its real value being burning competitors' money — human-labeled data is the biggest cost item (citing Mercor's revenue as evidence); he also conceded that IQ-level model gains don't mainly come from distillation, but argued the distillation strategy gives Chinese labs an edge in competition.
Why it matters
- The debate cuts straight to where capability gains in frontier models actually come from: if progress is driven mainly by RL environments rather than distillation, the key path to catching up with frontier models shifts from data/SFT to RL training infrastructure and environment building, directly affecting how we assess the US-China model gap and its closability.
- The discussion on how to obtain SFT trajectory data (hint-removal, sampling from open-source models) reveals how frontier labs reduce reliance on human annotation, with practical engineering value.
2026-09-25 ~ 2026-09-25 · 5 related posts
Primary sources
- [source] xeophon: distillation is irrelevant to why Chinese models are closing the gap — xeophon · 2026-09-25
- [source] Distillation debate: SFT era is over, RL environments drive capability gains now — xeophon · 2026-09-25
- Debate on sourcing SFT trajectories: prompt hints then remove, or use open models — fleetwood___ · 2026-09-25
- Distillation or frontier training? AI circle debates the real source of model gains — xeophon · 2026-09-25
- [source] Distillation is necessary but not sufficient, and it drains rivals' capital — fleetwood___ · 2026-09-25