DeepSeek V4 Flash Benchmarks: n_max=3 Yields 1.39× Speedup
Responsible_Pain3278 · reddit · 2026-08-18
The author conducted a week-long performance test of DeepSeek-V4-Flash-0731 on a Strix Halo (Ubuntu, 128GB) using Speculative Decoding.
- Setup: Model quantized to UD-IQ3XXS with Q6 attention. Compared DSpark Drafter Q80 vs. self-quantized Q2KS.
- Key Findings:
- Optimal nmax: nmax=3 is the peak, achieving 28.5 tok/s (1.39× over the 20.48 tok/s baseline).
- Quantization: Q2 and Q8 perform within 1-3% of each other; the smaller Q2 (6.45 GB) is preferred.
- Task Variance: Code/Math see significant gains (up to 1.52×), while Prose/Translation see modest gains (1.1-1.2×). nmax 5-7 only helps with high-acceptance tasks like repetition.
- Config: The author provides the final llama-server command (including ngram-mod, max thinking effort) and notes that while the sweep used a 64k window, real-world mixed usage at 128k should expect 22-28 tok/s.
More from Infra
- Volcengine OpenViking: Self-evolving context DB for agents — volcengine · 2026-08-18
- FreeToken framework claims major MoE inference speedups — wavefnx · 2026-08-18
- FreeToken Claims Faster MoE Inference vs. llama.cpp and Ollama — wavefnx · 2026-08-18
- Epoch AI: Musk Is the Only Frontier Lab CEO Building Data Centers — soumitrashukla9 · 2026-08-18
- AI server demand polarizes MLCC lead times, high-end hits 10 months — zephyr_z9 · 2026-08-18
- NVIDIA pledges up to $105B in project guarantees, raising questions about artificial growth fueled by self-funding chip sales — heypearlai · 2026-08-18