DeepSeek ships research artifacts, not products — explaining its eval gaps
teortaxesTex · x · 2026-09-11
Long-time observer teortaxesTex argues the key to understanding DeepSeek is that, unlike other labs, it ships internal research artifacts rather than polished products, with a policy of basically not caring about users. That explains high internal evals, low external scores, brittleness, and weird capability gaps.
Related event: DeepSeek Criticized as Research-Heavy, Product-Light(2 posts)→
More from Models
- DeepSeek's new open-source model reportedly crushes GLM and Kimi at 4-10x lower prices — anselm · 2026-09-11
- OpenAI launches GPT-Live-1 voice agent API, and running it 24/7 costs $72/day — gabrielchua · 2026-09-11
- GPT-6 Astra rebuilds Cessna 337 landing gear from a YouTube video; Fable 5.1 falls short — FinanceYF5 · 2026-09-11
- 'If Fable wasn't AGI, neither is Astra' — and the debate is collapsing into definitions — haider1 · 2026-09-11
- OpenAI's superintelligence reportedly tackling all Millennium Prize Problems — Dr_Singularity · 2026-09-11
- Users Suspect Anthropic's New Model Mythos Was Trained on CAPTCHAs — voooooogel · 2026-09-11