Open-Source LLM Distillation Debate: Why Don't Thinking Traces Yield Similar Gains?

yacineMTB · x · 2026-08-13

A discussion revolving around the technical details of large language model distillation. Some question that if thinking traces and chain-of-thought are widely believed to be effective for distillation, why haven't other models achieved similar performance boosts, especially given the internet is already flooded with massive amounts of open-source code generated by top-tier models. In response, the original commenter points out that true distillation might actually require access to the model's logits (probability distribution) rather than just the text output.

Original post →

More from Models

Models channel →