Steve Hsu: DeepSeek's efficiency comes from architecture, not distillation

jedisct1 · x · 2026-09-12

Steve Hsu argues that non-technical observers (and ideologically bent technical ones) overemphasize distillation while ignoring the model architecture innovations that only Chinese labs publish. A model like DeepSeek V4.1 owes its efficiency primarily to architecture.

He notes that while weights can be improved via distillation, the FLOP requirements for RL rollouts are large, and distilled traces can't be the main driver of model quality — DeepSeek has to get its own RL working well. He adds that DeepSeek is fairly open about its RL details, though the identity of its "teacher models" and data partners remains unclear.

Original post →

More from Models

Models channel →