llama.cpp adds DFlash speculative decoding for Qwen3.8-27B, faster than MTP

victormustar · x · 2026-10-06

llama.cpp author Georgi Gerganov announced that users running Qwen3.8-27B with MTP speculative decoding can switch to the new DFlash draft mode for extra speed: llama serve -hf ggml-org/Qwen3.8-27B-GGUF --spec-type draft-dflash --spec-draft-n-max 7, requiring llama.cpp v0.6.0. He contrasted it with the previous draft-mtp setup in a quoted post.

Original post →

More from Infra

Infra channel →