Distillation Debate: Frontier CoTs Too Off-Policy For Tiny Models, 27B Is The Better Teacher

JoshPurtell · x · 2026-09-03

In an X debate over whether a major lab distilled another model, JoshPurtell breaks distillation into two paths: (1) jailbreaking APIs to copy raw chain-of-thought — highly suspicious and likely to get accounts banned; (2) running the model on dev tasks and training on tool calls/non-reasoning outputs — far harder to detect.

His core claim: frontier reasoning traces distill poorly into tiny models like Qwen 3.5 0.8B because they're too off-policy (too smart/terse); a 27B teacher is the better fit. He adds that single-shot simple tasks like Shopify's don't benefit much from copying frontier outputs, while long-horizon and coding tasks might. Counterparty VivaLaPanda pushes back, noting that distilling a big general model into a smaller task model is standard practice across the industry.

Related event: JoshPurtell breaks down the model distillation debate: task type determines risk and feasibility(8 posts)→

Original post →

More from coding & agent

coding & agent channel →