DeepSeek-V4-Flash benchmarks: 286 tok/s single-stream, +34% boost with speculative decoding
mcraddock · x · 2026-08-19
Benchmark results for DeepSeek-V4-Flash on a single DGX Station GB300 have been shared. The model achieves 286 tok/s single-stream and 4,553 tok/s at 32 concurrent requests, with a WikiText-2 perplexity of 5.128. Notably, switching to dspark speculative decoding using 7 draft tokens yielded a 34% performance improvement over MTP-3.
More from Infra
- Samsung raises advanced chip foundry prices by up to 15% amid AI demand — firstadopter · 2026-08-19
- Samsung raises advanced chip prices up to 15% on tight AI demand — firstadopter · 2026-08-19
- HSWQ quantization outperforms native NVFP4 for image generation — Zestyclose_Bake3680 · 2026-08-19
- llama.cpp PR adds CPU offload for dense model FFN layers — pmttyji · 2026-08-19
- Ds4 v0.6.2: Runs DeepSeek 284B on single DGX Spark at 1000 tok/s — pbaylies · 2026-08-19
- opentel-mcp v0.11 fixes server-agent trace disconnection via W3C — Thirumalaiboobathi · 2026-08-19