Blogger challenges report DeepSeek trained R2 on Huawei Ascend, citing shaky details
teortaxesTex · x · 2026-09-25
Pushing back on reports that DeepSeek tried training an "R2" model on Huawei Ascend chips, teortaxesTex raises several objections:
- DeepSeek likely never had a separate "R2" project: its strategy has been to unify model lines and serve as few checkpoints as possible, with DS-R more likely to spin off like DS-Coder. Testing a flagship on a brand-new stack makes little sense versus a V3.5-lite trial.
- Supporting evidence: Zhipu trained a 9B GLM-Image on Ascend, yet still trains its flagships on Nvidia months later — even a state-aligned lab stayed cautious.
- Huawei's best hardware at the time (910C) had poor software support; the first large model trained on such hardware only recently came from Meituan — which is implausibly better at ML than DeepSeek.
Responding to "Will Huawei catch up to Nvidia by 2030?", he argues the post's assumptions deserve scrutiny: if China escalated HBM sourcing to a national effort, the conclusion could flip entirely.
Related event: Blogger disputes report that DeepSeek trained R2 on Huawei Ascend(2 posts)→
More from Infra
- Caching policy lookups per inode cuts eBPF security agent CPU cost ~90% — JeremyCMorgan · 2026-09-25
- Photonic matrix core on thin-film lithium niobate runs in-situ backpropagation at 8-bit precision — jwt0625 · 2026-09-25
- Report: Google to Launch TPUs Into Space Next Week Aboard Falcon 9 to Test Orbital AI Data Centers — VariationLivid3193 · 2026-09-25
- Opus 5.5 Tops SimpleBench at 88.4%, Xiaomi Open-Sources MiMo-V2.6-Pro Near Frontier — Latent Space · 2026-09-25
- Skip the PC case: open-air rigs are the way for 3+ RTX Pro 6000 setups — knowrohit07 · 2026-09-25
- DeepSeek's DSec sandbox platform serves 3M sandboxes daily for agent RL training; Liang Wenfeng co-authors — teortaxesTex · 2026-09-25