DeepSeek-V4-Flash benchmarks: 286 tok/s single-stream, +34% boost with speculative decoding

mcraddock · x · 2026-08-19

Benchmark results for DeepSeek-V4-Flash on a single DGX Station GB300 have been shared. The model achieves 286 tok/s single-stream and 4,553 tok/s at 32 concurrent requests, with a WikiText-2 perplexity of 5.128. Notably, switching to dspark speculative decoding using 7 draft tokens yielded a 34% performance improvement over MTP-3.

Original post →

More from Infra

Infra channel →