DeepSeek v4 runs at 47 tok/s on single GPU, passes 370k needle test

EAccelerate_42 · x · 2026-08-21

DeepSeek v4 Flash 0731 has been optimized to run smoothly on a single DGX Spark unit. Using EXL3 quantization, it achieves 47 tok/s generation speed and 1024 tok/s prefill, while passing a 370k token needle test. The performance is enabled by native NVFP4 KV cache, with quality comparable to Q4KM / Q5 GGUF.

Original post →

More from Infra

Infra channel →