Distillation Debate: Frontier CoTs Too Off-Policy For Tiny Models, 27B Is The Better Teacher
JoshPurtell · x · 2026-09-03
In an X debate over whether a major lab distilled another model, JoshPurtell breaks distillation into two paths: (1) jailbreaking APIs to copy raw chain-of-thought — highly suspicious and likely to get accounts banned; (2) running the model on dev tasks and training on tool calls/non-reasoning outputs — far harder to detect.
His core claim: frontier reasoning traces distill poorly into tiny models like Qwen 3.5 0.8B because they're too off-policy (too smart/terse); a 27B teacher is the better fit. He adds that single-shot simple tasks like Shopify's don't benefit much from copying frontier outputs, while long-horizon and coding tasks might. Counterparty VivaLaPanda pushes back, noting that distilling a big general model into a smaller task model is standard practice across the industry.
More from coding & agent
- Agno 3.0 ships with 100% ARC-AGI-3, persistent Python kernel and durable background runs — pritisinghhhh · 2026-09-03
- Hand Codex a Video URL and It Analyzes Weird Flowing-Water Harmonics — johnowhitaker · 2026-09-03
- Devnexus, largest US Java conference, retools 2027 around AI with 10 tracks — mkheck · 2026-09-03
- Auth0Lab opens up: first public look at how it does 0-to-1 products — yenkel · 2026-09-03
- Notch says AI coding turns a 10x engineer into 100x after Fable 5.1 flow session — majidmanzarpour · 2026-09-03
- Harvard Proposes Agentic Data Cracking, Cuts Unstructured QA Cost 53% on FanOutQA — Harvard · 2026-09-03