DeepSeek 'Bad at Post-Training'? Community Pushes Back Citing GRPO and Coding Models
teortaxesTex · x · 2026-08-13
Pushing back against rumors that DeepSeek struggles with post-training, a developer pointed out that the lab released the best open-weights coding model by October 2023, and that half the industry relies on their GRPO algorithm.
The original thread mentioned rumors that founder Liang Wenfeng deprioritized AI agents, allegedly leading to the departure of a key member. Commenters remained skeptical about their software engineering (SWE) agent capabilities, noting severe context-loss failures in the recent 0813 release.
Related event: Community Refutes Claims of DeepSeek's Weak Post-Training(2 posts)→
More from Companies & People
- Cognition Engineers Win DEF CON CTF with Perfect Blue Team — silasalberti · 2026-08-13
- The AI Race: A Marathon You Have to Sprint Through Entirely — peterwildeford · 2026-08-13
- Replit debuts at #24 on the Inc. 5000 list of fastest-growing companies — amasad · 2026-08-13
- KubeCon China Heads to Shanghai in Sept, Spotlighting AI Infra and Agents — PyTorch · 2026-08-13
- Paul Graham on Startups: 10% Weekly Growth and Catching the Second Wave — ycombinator · 2026-08-13
- Definitive Guide: Why You Can’t Copy Palantir — alexeyguzey · 2026-08-13