Distillation or frontier training? AI circle debates the real source of model gains
xeophon · x · 2026-09-25
A debate over where frontier model gains actually come from: one side argues the improvements are "largely" from distillation (and thus SFT), questioning whether the approach bottoms out—"some model" must actually be able to perform the task, or you have to buy trajectories from labellers.
xeophon pushes back, noting that argument doesn't hold: you can target different pass rates during training, and to advance the frontier a model has to be trained in a way that genuinely explores, not just distills existing capability. The exchange touches the core tension between distillation and self-exploration in RL for LLMs.
Related event: AI Researchers Debate Whether Distillation Drives China's Model Gains(5 posts)→
More from Models
- LightOn launches Ettin model suite: fine-tune a specialized classifier in 42 seconds — IgorCarron · 2026-09-26
- Perceptron Mk1.5 goes live with text, thinking, tools, and audio modes — rohanpaul_ai · 2026-09-26
- AI Detector Flags Enterprise Job Posting as 93% AI, Zero Substance — bushuev_online · 2026-09-26
- Should OpenAI Keep GPT-6 Astra's Successor Internal? The Compute Gap Problem — haider1 · 2026-09-26
- ChatGPT drops text chat limits for free users and upgrades default model — Aiden_Tech_Ai · 2026-09-26
- New Gemini 4 Pro checkpoint spotted in Arena, outranking GPT-6 Astra — cedric_chee · 2026-09-26