DFlash2 speculative decoding hits 2.8x speedup in local Qwen3.8 three-way benchmark

FantasticNature7590 · reddit · 2026-10-03

The author benchmarked three local Qwen3.8 builds — RadixArk 27B NVFP4 (dense), orcarouter 27B Uncensored, and RadixArk Flash-Next NVFP4 (MoE) — on identical 10 tests on a single RTX PRO 6000 (96GB), using SGLang v0.5.20 and vLLM v0.29.0.

Key findings:

Original post →

More from Infra

Infra channel →