incoai releases Qwen3.8-27B-DFlash2 draft model for speculative decoding
incoai · hf · 2026-08-20
incoai released Qwen3.8-27B-DFlash2, a block-diffusion based DFlash2 draft model for speculative decoding to speed up inference. It supports SGLang and vLLM and is trending on Hugging Face.
Related event: Qwen3.8-27B DFlash2 Draft Model Speeds Up Inference via Block-Diffusion(2 posts)→
More from Models
- Gemini 3.7 Flash solves complex network configs — DynamicWebPaige · 2026-08-20
- Qwen3.8-27B Test: Lower KV Cache Quantization Impacts Reasoning Quality — fbms2 · 2026-08-20
- GPT-5.6 learns to look up information online during evals — dejavucoder · 2026-08-20
- DeepSeek V4 Flash 0731 Pricing and Specs Revealed — AccBalanced · 2026-08-20
- Superwhisper releases S1-mini: 0.6B parameter text normalizer for STT output — Recoil42 · 2026-08-20
- Zhipu GLM 5.3 Demo: Task Execution via Pure Commands — BLUECOW009 · 2026-08-20