Training mixtures are now all synthetic: small-model training is really distillation

RexDouglass · x · 2026-09-07

Gregory Diamos argues that every source in modern training mixtures is a curated artifact built with large models, making training models at this size essentially distillation. He adds that when models of this size were last studied seriously, such corpora did not exist — so old findings about small-model capabilities may not carry over to the synthetic-data era.

Original post →

More from Models

Models channel →