Qwen3.6-27B speculative decoding benchmark finds DFlash fastest, with up to 4.6x speedup

thavoc77 · reddit · 2026-07-21

Benchmarks compare speculative decoding methods on Qwen3.6-27B (dense, NVFP4) on a single RTX PRO 6000 Max-Q, using the same client and three restart samples per point across vLLM and SGLang.

Main results

Engineering notes

Original post →

More from Infra

Infra channel →