View: DeepSeek should focus on post-training rather than 4.1

bookwormengr · x · 2026-08-18

The author argues that DeepSeek should not focus on version 4.1 but instead aim for a larger pre-training run. It notes that Moonshot gained an advantage by being the first to launch a 3T-tier model and can now focus on post-training. The comment also highlights the significant performance gains possible through post-training, citing the progression from GLM 5 to 5.3.

Related event: Opinion: DeepSeek Should Skip 4.1 and Go Bigger on Pretraining(2 posts)→

Original post →

More from Companies & People

Companies & People channel →