Qwen3.8-27B gets DFlash2 speculative-decoding GGUF release for llama.cpp

incoai · hf · 2026-08-22

Hugging Face user incoai released Qwen3.8-27B-DFlash2-GGUF, now trending. It is a quantized build of Qwen/Qwen3.8-27B featuring a DFlash2 draft model with speculative decoding, designed for llama.cpp-based local inference to speed up text generation on consumer hardware. Apache-2.0 licensed.

Original post →

More from Infra

Infra channel →