Steve Hsu: DeepSeek's efficiency comes from architecture, not distillation
jedisct1 · x · 2026-09-12
Steve Hsu argues that non-technical observers (and ideologically bent technical ones) overemphasize distillation while ignoring the model architecture innovations that only Chinese labs publish. A model like DeepSeek V4.1 owes its efficiency primarily to architecture.
He notes that while weights can be improved via distillation, the FLOP requirements for RL rollouts are large, and distilled traces can't be the main driver of model quality — DeepSeek has to get its own RL working well. He adds that DeepSeek is fairly open about its RL details, though the identity of its "teacher models" and data partners remains unclear.
More from Models
- Dev Claims DeepSeek v4.1 Flash "Way Better" Than Gemini Flash Models — gaganghotra_ · 2026-09-12
- LinkedIn user claims GPT-6 built a pixel-perfect Figma design system in 3 hours — AIandDesign · 2026-09-12
- Google reportedly achieves recursive self-improvement, new model due Oct 5 — Dr_Singularity · 2026-09-12
- Report: Google Achieved Recursive Self-Improvement in Next-Gen AI Model, Launching Oct 5 — Dr_Singularity · 2026-09-12
- Leaked Codenames Suggest OpenAI Model Tiers Map to Claude's Opus/Sonnet/Haiku — haider1 · 2026-09-12
- Eric Xing on "open source": open weights is a house with no blueprints — YiMaTweets · 2026-09-12