Which speculative decoding setup for local Qwen models? Two llama.cpp configs compared
cradlemann · reddit · 2026-09-21
Running Qwen3.8-Flash-Next-UD-Q4KXL and Qwen3.8-27B-UD-Q5KXL on an AMD Ryzen AI 9 HX 470 with llama-server's Vulkan backend, the author shares two speculative decoding configs: draft-mtp + ngram-map-k for the Flash model, draft-dflash + ngram-map-k for the 27B (spec-draft-n-max: 4, p-min: 0.5). They see little difference from the ngram decoders and note llama-bench can't benchmark speculative models without bias, asking what others use.
More from Infra
- Mozilla AI runs a local 30B model end-to-end to open a real bugfix PR, fully offline — mozilla-ai · 2026-09-21
- Cohere Labs launches Local AI community program for local inference and hardware tuning — Cohere_Labs · 2026-09-21
- Gewell: custom Gemma 4 inference engine cuts KV cache VRAM to 0.625x, losslessly — stoppableDissolution · 2026-09-21
- DeepSeek-V4.1-Flash redesigns the Transformer for agents, cutting KV cache to 890 bytes/token — AndLukyane · 2026-09-21
- Jev Engineering gives agents a decision brain, 193x faster and 444x cheaper in tests — agihouse_org · 2026-09-21
- UK's £225m Isambard-AI supercomputer cost about the same as one road bridge — charlieharris01 · 2026-09-21