Qwen3.8-27B GGUF Release with Speculative Decoding Support
z-lab · hf · 2026-08-23
z-lab released a GGUF quantized version of Qwen3.8-27B on Hugging Face. Based on the original Qwen model, it integrates DFlash2 and speculative decoding techniques using a draft model to accelerate inference, optimized for llama.cpp compatibility.
Related event: Qwen3.8-27B GGUF quantized release adds speculative decoding speedup(2 posts)→
More from Models
- ox model reviewed: meticulous PhD janitor as a long-horizon subagent — teortaxesTex · 2026-08-23
- MiniMax H3's high prompt adherence creates stiff, frozen videos lacking subtle motion — DifficultAd5938 · 2026-08-23
- Princeton's i1: A fully open text-to-image model backed by 300 controlled experiments — 机器之心 · 2026-08-23
- AI models exhibit 'Fablish' writing quirks: garbled negations and OSV word order — alexisgallagher · 2026-08-23
- Hands-on: Terra beats Sonnet, Flash unmatched on speed and quality — cgarciae88 · 2026-08-23
- Users Report Gemini Answering Coding Questions With Totally Unrelated Australian Tax Info — mic_n · 2026-08-23