Benchmark: DFlash2 vs. MTP speculating decoding on Qwen 3.8 27B

Opening-Broccoli9190 · reddit · 2026-08-24

A llama.cpp benchmark comparing DFlash2 and MTP speculative decoding techniques on the Qwen 3.8 27B model using an RTX 5090. Results show DFlash2 Q4 offers the highest speed (154 tps), while Q2 maximizes context size. MTP provides the largest context but lags in the speed-context balance metric. Q8 quantization is recommended against due to high memory usage with no performance benefit.

Related event: DFlash2 Shows Superior Performance in Qwen 3.8 27B Tests(2 posts)→

Original post →

More from Infra

Infra channel →