Training mixtures are now all synthetic: small-model training is really distillation
RexDouglass · x · 2026-09-07
Gregory Diamos argues that every source in modern training mixtures is a curated artifact built with large models, making training models at this size essentially distillation. He adds that when models of this size were last studied seriously, such corpora did not exist — so old findings about small-model capabilities may not carry over to the synthetic-data era.
More from Models
- 'Not AGI': User Spent $200 in 8 Hours on Astra, Found 3 of 4 Tasks Broken — Firm-Club-8334 · 2026-09-07
- Similarweb: ChatGPT's AI traffic share falls from 73.3% to 55.5% in 12 months — gaganghotra_ · 2026-09-07
- GPT-6 Astra (and Pro?) spotted on Simple-Bench leaderboard — From_Internets · 2026-09-07
- Testing Gemini as music understanders: Pro 3.1 solid, Flash models hallucinate sounds — teropa · 2026-09-07
- GPT-6 Astra posts 65.6% on ClockBench — Well_being1 · 2026-09-07
- Tencent Hunyuan ships Hy4 preview upgrade: same quality, fewer turns, lower token usage — xeophon · 2026-09-07