z-lab's Qwen3.8-27B DFlash2: block-diffusion draft model for fast inference

z-lab · hf · 2026-08-19

z-lab released Qwen3.8-27B-DFlash2, a block-diffusion draft model for speculative decoding acceleration, compatible with sglang and vLLM. It's trending on Hugging Face.

Original post →

More from Infra

Infra channel →