View: DeepSeek should focus on post-training rather than 4.1
bookwormengr · x · 2026-08-18
The author argues that DeepSeek should not focus on version 4.1 but instead aim for a larger pre-training run. It notes that Moonshot gained an advantage by being the first to launch a 3T-tier model and can now focus on post-training. The comment also highlights the significant performance gains possible through post-training, citing the progression from GLM 5 to 5.3.
Related event: Opinion: DeepSeek Should Skip 4.1 and Go Bigger on Pretraining(2 posts)→
More from Companies & People
- Theo and Team Launch 'MEGA' Course for AI-Native Engineering — johnlindquist · 2026-08-18
- Stop posting the misleading ChatGPT growth graph, data comes from app merge — zephyr_z9 · 2026-08-18
- Critique: Modern Startups Are Soulless and 'TikTokified' — moonsandhues · 2026-08-18
- Users list fastest falling AI companies, noting Perplexity and Stability AI stalls — Angaisb_ · 2026-08-18
- Gen Z Dominates AI Engineering Hires with 69% Share in 2025 — rohanpaul_ai · 2026-08-18
- SF + LA Tech Week Call for Proposals Closes This Friday — andrewchen · 2026-08-18