Qwen3.8-27B gets DFlash2 speculative-decoding GGUF release for llama.cpp
incoai · hf · 2026-08-22
Hugging Face user incoai released Qwen3.8-27B-DFlash2-GGUF, now trending. It is a quantized build of Qwen/Qwen3.8-27B featuring a DFlash2 draft model with speculative decoding, designed for llama.cpp-based local inference to speed up text generation on consumer hardware. Apache-2.0 licensed.
More from Infra
- AgenticROS adds Organizations and Teams support for robot sharing — chrismatthieu · 2026-08-22
- New narrative for AI datacenters: Jobs, lower taxes, and better infrastructure — alexvoica · 2026-08-22
- Ramp launches Router LLM gateway to cut inference costs by 40% — round · 2026-08-22
- MCP Isn't Replacing APIs: It's Changing Who APIs Are Designed For — kush_patil · 2026-08-22
- Data Center Opposition Surged from 42 to 75 Percent in One Year — The Decoder · 2026-08-22
- Woolly post-trains Qwen3-8B for 2–3× faster math & code decoding — bosmeny · 2026-08-22