Schulman cites model trained only on pre-1930 text beating Claude 3 Opus after distillation

victor_explore · x · 2026-09-13

OpenAI co-founder John Schulman shared a distillation case: a model that had only read pre-1930 text and never seen code was fine-tuned on AI-generated coding examples and beat Claude 3 Opus on a coding benchmark. Baseten's Charlie O'Neill outlined where the approach fails. victorexplore adds the key variable: the teacher-student gap — cheap copying works when the gap is small and collapses when it's wide, so spend should go into building distillation rungs rather than one big fine-tune.

Original post →

More from Models

Models channel →