Why do Sonnet 5 and Opus 5 feel worse than GLM 5.3? Distillation's limits, dissected

baseten · x · 2026-09-12

Dwarkesh Patel highlights a podcast segment where John, Beren, and Charlie speculate why Sonnet 5 and Opus 5 feel like worse models than GLM 5.3 — even though Anthropic could do raw logit distillation from Fable and train on the same environments. The discussion covers the value of distillation, what effective distillation takes, and which model behaviors resist being extracted via distillation.

Original post →

More from Models

Models channel →