DeepSeek's V3-to-V4 gap blamed on failed architecture exploration, Ascend training rumor denied
teortaxesTex · x · 2026-09-25
Commenting on the release gap between DeepSeek V3 and V4, teortaxesTex argues it is explained by a series of architecture explorations that failed at scale, starting with NSA (sparse attention), and speculates experiments on Ascend informed the eventual focus on TileLang.
He also flatly denies the Epoch-related rumor: neither DeepSeek R2 nor "training a major product model on Ascends" ever happened — calling it a broken-telephone rumor that only reflects on Epoch's evidentiary standards.
More from Models
- Dev Compares Codex vs Claude: Claude Nails Multi-Agent Orchestration, Codex Doesn't — madhavsinghal_ · 2026-09-25
- Student Pays $30/Month for Google AI Pro, Still Finds Gemini Unreliable and Hallucinatory — makeitbumthem · 2026-09-25
- Meta's Muse app hit 1.8M iOS downloads in 12 days, beating ChatGPT's 1.3M — Beth_Kindig · 2026-09-25
- Opus 5.5's token usage is far more reasonable: 13-hour session on a 20x Max plan — majidmanzarpour · 2026-09-25
- TypeSafe's Jev: A Model That Can't Write but Decides, Claiming 200x Speed and 400x Cost Savings — rseroter · 2026-09-25
- Prime Intellect pitches highest-throughput GLM 5.3 inference with OpenAI-compatible eval API — willcb · 2026-09-25