DeepSeek 'Bad at Post-Training'? Community Pushes Back Citing GRPO and Coding Models

teortaxesTex · x · 2026-08-13

Pushing back against rumors that DeepSeek struggles with post-training, a developer pointed out that the lab released the best open-weights coding model by October 2023, and that half the industry relies on their GRPO algorithm.

The original thread mentioned rumors that founder Liang Wenfeng deprioritized AI agents, allegedly leading to the departure of a key member. Commenters remained skeptical about their software engineering (SWE) agent capabilities, noting severe context-loss failures in the recent 0813 release.

Related event: Community Refutes Claims of DeepSeek's Weak Post-Training(2 posts)→

Original post →

More from Companies & People

Companies & People channel →