DeepSeek-V4-Flash tested: hits 2,100 tok/s aggregate on 4x RTX PRO 6000
BanghuaZ · x · 2026-08-01
Shared benchmark results for the rumored DeepSeek-V4-Flash-0731 model running on 4x RTX PRO 6000 (SM120).
- Stack: Deployed via SGLang and DSpark.
- Throughput: 200–225 tok/s single-stream; 2,100 tok/s aggregate at 32 concurrent requests.
- Accuracy & Fix: Achieves 0.99 on GSM8K. Required a 3-line FlashInfer patch (topk=192) to run on the new architecture.
The author notes that all figures are measured and fully reproducible.
More from Infra
- Laguna Doubles Performance: Significant Mac Inference Speedup Without Speculative Decoding — gajesh · 2026-08-01
- Will the AI Agent Explosion Overload and Break Internet Infrastructure? — Ok-Video4323 · 2026-08-01
- Kimi K3 Hits Record 172 Tokens/sec in Inference Speed — AccBalanced · 2026-08-01
- OpenAI Slashes GPT-5.6 Prices, MiniMax Launches H3, ByteDance Unveils Seedance 2.5 — 创业邦 · 2026-08-01
- SpaceX is Hiring to Build the Most Powerful AI Supercomputer Clusters On and Off Earth — elonmusk · 2026-08-01
- audio.cpp 0.5 Released: Dramatic TTS and Cross-lingual Voice Transfer — Acceptable-Cycle4645 · 2026-08-01