DFlash 2 speeds up GLM-5.3-Flash by 2.8x via block-diffusion speculative decoding

Lianhuiq · x · 2026-08-28

Inco AI released the DFlash 2 draft model for GLM-5.3-Flash, achieving a 2.8x speedup using speculative decoding. DFlash 2 employs a block-diffusion mechanism to predict a whole block of tokens in one pass and track coherent paths with a lightweight selector. Benchmarks run on SGLang with NVIDIA GB300 GPUs show lossless decoding where greedy outputs match the target model exactly.

Related event: Inco AI Releases DFlash 2 Draft Model, Speeding Up GLM-5.3-Flash by 2.8x(3 posts)→

Original post →

More from Infra

Infra channel →