DFlash2 speculative decoding hits 2.8x speedup in local Qwen3.8 three-way benchmark
FantasticNature7590 · reddit · 2026-10-03
The author benchmarked three local Qwen3.8 builds — RadixArk 27B NVFP4 (dense), orcarouter 27B Uncensored, and RadixArk Flash-Next NVFP4 (MoE) — on identical 10 tests on a single RTX PRO 6000 (96GB), using SGLang v0.5.20 and vLLM v0.29.0.
Key findings:
- Speed: attaching the DFlash2 drafter to the 27B lifted Spec-Bench from 75 to 210 tok/s (2.8x) for one user; DFlash2 won on both engines (3.7 drafted tokens/step vs 2.9 for MTP). Full-window prefill: Flash-Next 22.4s vs 97s for the 27B. BF16 runs at just 97 tok/s, showing large quantization gains.
- Engines: SGLang was faster overall but mostly due to the checkpoint; the same export ran 16% faster on vLLM (169 vs 146 tok/s).
- Capability: Flash-Next won 5 of 10 tests, the 27B won 4. Tool use (BFCL subset): 27B 73.3%, Uncensored 70.8%, Flash-Next 64.5%. In the battle arena the 27B scored 700/1000, beating Claude Fable 5.1 (678) and GPT-5.6 (473). Only Flash-Next completed the Rube Goldberg task; both 27Bs burned their 111K-token budgets on thinking.
- 4-user concurrent throughput with RadixArk NVFP4 + DFlash2 reached 607 tok/s.
More from Infra
- Burkov: Generative AI Only Makes Money for GPU Sellers, Echoing Dotcom Bubble — burkov · 2026-10-03
- Deriving KV-cache placement from abstract representations: prefill and inference are linked — vtabbott_ · 2026-10-03
- NVIDIA long stopped just selling GPUs: from CUDA to sovereign, agentic and physical AI — sudoraohacker · 2026-10-03
- Debate: Nvidia's Moat Is Its Software Stack, Not Just Hardware Lock-In — QuintinPope5 · 2026-10-03
- Cerebras CEO Explains Why Wafer-Scale SRAM Beats GPU HBM by 2500x in LLM Inference — rohanpaul_ai · 2026-10-03
- Cerebras CEO Flexes His Own 42MW 13.8kV Power Generator for AI Compute — dunkhippo33 · 2026-10-03