DFlash 2 Introduces Block-Diffusion Speculative Decoding to Speed Up GLM-5.3-Flash

songhan_mit · x · 2026-08-28

Inco AI released the DFlash 2 draft model, utilizing block-diffusion technology for speculative decoding to accelerate the inference speed of zai-org/GLM-5.3-Flash.

Related event: Inco AI Releases DFlash 2 Draft Model to Speed Up GLM-5.3-Flash Inference(2 posts)→

Original post →

More from Infra

Infra channel →