llama.cpp adds DFlash speculative decoding for Qwen3.8-27B, faster than MTP
victormustar · x · 2026-10-06
llama.cpp author Georgi Gerganov announced that users running Qwen3.8-27B with MTP speculative decoding can switch to the new DFlash draft mode for extra speed: llama serve -hf ggml-org/Qwen3.8-27B-GGUF --spec-type draft-dflash --spec-draft-n-max 7, requiring llama.cpp v0.6.0. He contrasted it with the previous draft-mtp setup in a quoted post.
More from Infra
- Ex-UK energy official warns AI needs 500GW by 2035 and the industry isn't ready — ShakeelHashim · 2026-10-06
- Starlink now has 11,000+ satellites in orbit, two-thirds of all active satellites — XFreeze · 2026-10-06
- Ben Bajarin: Agentic AI Will Drive Datacenter CPU Demand, Scale-Up Domain Is the Key Battleground — BenBajarin · 2026-10-06
- GLM 5.3 full NVFP4 deployable on 4x B200 or H200 with Marlin kernels — TheZachMueller · 2026-10-06
- Dev quantizes GLM-5.3-UNCENSORED to MXFP4, cutting size 44% for AMD GPUs — bakatristan · 2026-10-06
- OpenAI reportedly spent tens of millions in compute to crack Navier-Stokes in days — haider1 · 2026-10-06