llama.cpp Benchmark: DFlash2 20% Faster but Cuts Context by 38%

Opening-Broccoli9190 · reddit · 2026-08-21

A user benchmarked DFlash2 vs MTP speculation decoding strategies in llama.cpp using Qwen 3.8 27B (Q8) on a 5090RTX.

Key Data:

Conclusion:

Full llama-server configuration parameters are included.

Related event: llama.cpp's New DFlash2 Boosts Qwen Inference Speed Up to 3x(2 posts)→

Original post →

More from Infra

Infra channel →