Distilling big-RL models into small ones breeds overconfident agents without calibration RL

willcb · x · 2026-10-05

willcb argues that the common "big model RL -> smaller model distill" pipeline has a structural flaw: a general-purpose smaller distill without follow-up RL for calibration will necessarily be overconfident and under-use reasoning, because awareness of one's own capabilities isn't carried over by distillation. Conclusion: distillation shouldn't be the endpoint—small models need additional RL to recalibrate their sense of their own limits.

Original post →

More from Models

Models channel →