Qwen3.8-27B DFlash2 Draft Model Speeds Up Inference via Block-Diffusion

z-lab and incoai released Qwen3.8-27B-DFlash2, a block-diffusion-based draft model for speculative decoding that accelerates inference and is compatible with frameworks like SGLang.

2026-08-19 ~ 2026-08-20 · 2 related posts