DeepSeek v4 runs at 47 tok/s on single GPU, passes 370k needle test
EAccelerate_42 · x · 2026-08-21
DeepSeek v4 Flash 0731 has been optimized to run smoothly on a single DGX Spark unit. Using EXL3 quantization, it achieves 47 tok/s generation speed and 1024 tok/s prefill, while passing a 370k token needle test. The performance is enabled by native NVFP4 KV cache, with quality comparable to Q4KM / Q5 GGUF.
More from Infra
- Liquid AI's DSpark drafters hit 2x decode speed on iPhone within 24 hours — helloiamleonie · 2026-08-21
- LLMRouter: Open-source library balances cost and quality with 16 routing methods — adnan_hashmi · 2026-08-21
- Nvidia Reportedly in Talks to Invest in Data Center Power — Polymarket · 2026-08-21
- OpenAI's 8-GW Campus: 35K Construction Jobs vs. 2.5K Operating Jobs — Crescitaly · 2026-08-21
- OpenAI monitoring adds 20% inference overhead, reshaping agent economics — Crescitaly · 2026-08-21
- Auditing safety signals in Zero Data Retention without provider access — Crescitaly · 2026-08-21