DeepSeek’s next model should add continual learning before chasing bigger gains
teortaxesTex · x · 2026-07-23
The thread argues that DeepSeek’s next generation should prioritize continual learning. It points to the V4 paper and hiring signals as evidence that the team is moving toward that direction.
A reply adds that V4 should also have native multimodality, while another commenter suggests V4 is already at the largest scale DeepSeek can train right now. The implication is that the company may need to keep improving cost/performance before attempting the bigger leap.
Related event: DeepSeek Founder's Investor Call Reveals AGI Roadmap and Compute Plans(37 posts)→
More from Models
- DeepSeek V4.1 Flash Hits 98% of GPT-6 Astra's Score at 1.4% of the Cost in Third-Party Benchmark — ayushtweetshere · 2026-09-11
- TheZvi Polls: Has Your Coding Model Choice Changed Since Fable 5.1 and Astra? — TheZvi · 2026-09-11
- antirez Weighs In on Anthropic Banning Minors From Using Claude — antirez · 2026-09-11
- Engram's random reads don't suit SSDs; CPU-memory over NVLink could serve all 72 GPUs — bookwormengr · 2026-09-11
- Fully local voice assistant on an RTX 3060 replicates the GPT Live demo in 6.5 minutes — liampetti · 2026-09-11
- Meta's Muse Agent has built-in invite code logic, hinting at free-usage expansion — testingcatalog · 2026-09-11