Qwen3.8-27B GGUF Release with Speculative Decoding Support

z-lab · hf · 2026-08-23

z-lab released a GGUF quantized version of Qwen3.8-27B on Hugging Face. Based on the original Qwen model, it integrates DFlash2 and speculative decoding techniques using a draft model to accelerate inference, optimized for llama.cpp compatibility.

Related event: Qwen3.8-27B GGUF quantized release adds speculative decoding speedup(2 posts)→

Original post →

More from Models

Models channel →