Tech Debate: Can Models Be Distilled Using Only API Outputs?
tomekkorbak · x · 2026-07-20
The AI tech community recently debated the exact definition of model distillation on X.
- Core Controversy: Some techies pointed out that "distillation" is now loosely used to mean training solely via API outputs (generated text samples) without needing raw logits data.
- Technical Rebuttal: Developer @migtissera strongly questioned this, arguing that true distillation is impossible using only API outputs. He used this to refute recent accusations that Chinese AI labs boosted their models by "distilling API data," noting that companies like Anthropic even hide reasoning traces server-side, returning only hashes to make reverse extraction harder.
- Community Call to Action: Researcher Ryan Greenblatt called for a definitive article outlining the actual effects of "training on outputs only" and standardizing the use of the term "distillation" when logits are unavailable.
Related event: AI Model Distillation Debate: Normal Tech Evolution or IP Theft?(5 posts)→
More from Research
- GameWorld wins Best Paper Runner-Up at ECCV 2026 Multimodal Digital Agents Workshop — MikeShou1 · 2026-09-11
- 3D ResNet Paper Crosses 3,000 Citations Eight Years After CVPR 2018 — HirokatuKataoka · 2026-09-11
- Sample selection and ordering matter a lot in LLM training: DataFlex makes data scheduling dynamic — Puzzleheaded_Box2842 · 2026-09-11
- Jeff Heaton's Intro to the Math of Neural Networks eBook Is Free to Download — blaizedsouza · 2026-09-11
- Mathematician Daniel Litt Launches Problem Repo to Track Human vs AI Progress: 15 Problems, 1 Solved — littmath · 2026-09-11
- Open ECDSA.fail challenge uses AI agents to shrink Shor's-algorithm quantum circuits for Bitcoin keys — StefanoGogioso · 2026-09-11