DeepSeek Hits 243 tok/s on Dual RTX 6000 with Speculative Decoding
TheZachMueller · x · 2026-08-02
A developer tested DeepSeek-V4-Flash-0731 with DSpark speculative decoding on 2× RTX PRO 6000 Blackwell (TP=2).
- Performance: Achieved a median single-stream throughput of 243 tok/s, a 3.1× speedup over the no-spec baseline. Aggregate throughput hit 299 tok/s (c2) and 403 tok/s (c4).
- Stability: Draft acceptance rate was 74–76%, sustaining 237–245 tok/s over a 3-hour mixed-load soak test with zero failures.
- Bottleneck: Avoided the expected SM120 kernel wall. The real issue is a deterministic failure when maxnumseqs > 4 due to a hardcoded prefill chunk size feeding an empty slice to the sparse kernel. Capping it at 4 resolved the issue, making it the author's new daily driver.
More from Infra
- Multi-GPU Full VRAM Deployment of DeepSeek V4 Yields Only 600 t/s PP — fragment_me · 2026-08-02
- Report: SpaceX Plans 4GW Compute Cluster, Rivaling US Grid Expansion — BenBajarin · 2026-08-02
- China's AI Compute Pairings: DeepSeek x Huawei, Moonshot x Alibaba — bronzeagepapi · 2026-08-02
- Chamath Shares AI Investing Guide: Land and Power Offer Fastest Returns — davidyin44 · 2026-08-02
- Hyperscalers Report Accelerating Cloud Revenue: Google Cloud Up 82% — ai · 2026-08-02
- YC-backed Stoa launches GPU RFQ marketplace, sees $300M+ demand in first month — ycombinator · 2026-08-02