Why do Sonnet 5 and Opus 5 feel worse than GLM 5.3? Distillation's limits, dissected
baseten · x · 2026-09-12
Dwarkesh Patel highlights a podcast segment where John, Beren, and Charlie speculate why Sonnet 5 and Opus 5 feel like worse models than GLM 5.3 — even though Anthropic could do raw logit distillation from Fable and train on the same environments. The discussion covers the value of distillation, what effective distillation takes, and which model behaviors resist being extracted via distillation.
More from Models
- Continued pretraining vs RAG: an accuracy and performance comparison on Qwen 3.5 4B — funJS · 2026-09-12
- DeepSeek App Quietly Adds Four Read-Aloud Voices, Fueling New TTS Model Rumors — testingcatalog · 2026-09-12
- Watch OpenAI's Astra effortlessly drive a browser, as users ask about phone use next — infoxiao · 2026-09-12
- Researcher speculates spatial reasoning leap comes from Blender training data — yoavartzi · 2026-09-12
- Sakana's Fugu Max hits OpenRouter: multi-agent orchestration, 1M context at $2/$6 per 1M tokens — SakanaAILabs · 2026-09-12
- "Why did you nerf Astra?" Users report model degraded days after launch — Significant-Ad6970 · 2026-09-12