Schulman cites model trained only on pre-1930 text beating Claude 3 Opus after distillation
victor_explore · x · 2026-09-13
OpenAI co-founder John Schulman shared a distillation case: a model that had only read pre-1930 text and never seen code was fine-tuned on AI-generated coding examples and beat Claude 3 Opus on a coding benchmark. Baseten's Charlie O'Neill outlined where the approach fails. victorexplore adds the key variable: the teacher-student gap — cheap copying works when the gap is small and collapses when it's wide, so spend should go into building distillation rungs rather than one big fine-tune.
More from Models
- kalomaze: Opus 5 is "such a bad model" — kalomaze · 2026-09-13
- Wenhu Chen can't even understand many questions in AA-intelligence AI benchmarks — WenhuChen · 2026-09-13
- Commenter: The OpenAI Navier-Stokes proof would be hailed as a breakthrough if posted anonymously — skdh · 2026-09-13
- GPT-6 Astra skips the chat box: Plus users get just 5-45 messages per 5 hours, locked to Work and Codex — 新智元 · 2026-09-13
- LLM Architecture Gallery: one chart to compare every major LLM design — zainhas · 2026-09-13
- DeepSeek-V4.1-Flash reportedly live on Ollama cloud with zero data retention and off-peak pricing — johnseach · 2026-09-13