Draft model hits ~60 tok/s running Qwen3.8-27B at 131k context on a 16GB GPU
pneuny · reddit · 2026-09-12
- Using the DFlash2 draft model (HermiHg/Qwen3.8-27B-DFlash2-Q2KS-MIX-GGUF) with an IQ3XXS main model at 131k context, the author averaged 60 tokens/s on a 16GB RX 9070 XT.
- It beats the built-in MTP speculative decoding, which multiplies VRAM requirements; for single-thread mode this is among the best options for a 16GB GPU.
- Caveat: speculative decoding is more sensitive to VRAM overflow — if you exceed available VRAM, turning it off is actually faster on his DDR5 PCIe 5 system.
- More extensive testing on different cards is documented in the HF discussion thread.
More from Infra
- Glass core substrates show 2x better warpage than organic core without stiffener — jwt0625 · 2026-09-12
- d-Matrix partners with NVIDIA to plug Raptor XPUs into NVLink Fusion rackscale systems — bookwormengr · 2026-09-12
- Running out of context on a large codebase: how to auto-handoff long-running local LLM tasks — Developer-Y · 2026-09-12
- The Economist: Nvidia is the central bank of AI — tolugenius · 2026-09-12
- What MoE/LLM runs well offline on a 24GB M5 MacBook Air? — itis_whatit-is · 2026-09-12
- DeepSeek v4.1 Flash on-device test: q2 runs at 16 tok/s but tool calls go off the rails — challis88ocarina · 2026-09-12