DeepSeek V4.1-Flash hits 5,800 TPS per Ascend 950DT card, but critics call throughput underwhelming
teortaxesTex · x · 2026-10-04
- A developer reports deploying DeepSeek-V4.1-Flash with mixed FP8-FP4 precision on 32 Huawei Ascend 950DT cards: single-card decoding throughput reaches 5,800 TPS at 13.79 ms for 128K-sequence scenarios, while a low-latency config keeps latency under 5 ms at 2,727 TPS.
- teortaxesTex argues the results are weak: the throughput is only about one-third of what DeepSeek reportedly achieved per GPU with V4-Flash in the DSpark report, despite TileLang kernels. He suggests a larger deployment unit or future optimizations (e.g. npugraphex) may be needed.
More from Infra
- China's aggressive AI push runs into a new problem: too much usage, NYT reports — TMWNN · 2026-10-04
- UBS Lead Time Map: Foundries Take 3-4 Years While GPUs Ship in 6-12 Months — AccBalanced · 2026-10-04
- NVIDIA NIM API unreliable despite decent speed, users seek alternatives — Professional_Log1367 · 2026-10-04
- LLMxRay: Open-Source Local Observability for LLM Traffic, Launches With One Command — GuruCsharp · 2026-10-04
- Datacenter buildout would have been less aggressive in a world where everyone gets AGI at once — willcb · 2026-10-04
- UBS Sees Global Rack Capacity Hitting 104.5 GW by 2030, Half Going to Nvidia — AccBalanced · 2026-10-04