Huawei’s Ascend 950DT could beat Nvidia B300 on tokens per watt, analysis says
teortaxesTex · x · 2026-07-21
A quoted analysis argues that Huawei’s next-gen Ascend 950DT may beat Nvidia’s B300 on tokens-per-watt, while the current CloudMatrix 384 generation is already at parity with the token rates API providers target.
The comparison is based on announced specs, assumed power numbers, and calibration from Huawei’s published DeepSeek-R1 serving results, with K3 geometry inferred ahead of the 7/27 weight drop.
More from Infra
- 12 KV Cache Reduction Techniques Every AI Engineer Should Understand, Explained — blaizedsouza · 2026-09-11
- The shadow GPU capacity market is formalizing, with Meta selling excess compute to outside buyers — DavidLinthicum · 2026-09-11
- Engram's random reads don't suit SSDs; CPU-memory over NVLink could serve all 72 GPUs — bookwormengr · 2026-09-11
- 80% of the DIY LLM inference hype posters have already quit — it's brutally hard systems work — abhijithneil · 2026-09-11
- Hugging Face's Ultra Scale Playbook: a free book on training LLMs on GPU clusters — mdancho84 · 2026-09-11
- Is inference latency becoming the biggest bottleneck for production AI agents? — Euphoric_Sea632 · 2026-09-11