Inco AI Releases DFlash 2 Draft Model to Speed Up GLM-5.3-Flash Inference
Inco AI has released DFlash 2, a draft model that uses block-diffusion-based speculative decoding to accelerate zai-org/GLM-5.3-Flash inference by 2.8x.
2026-08-28 ~ 2026-08-28 · 2 related posts
- DFlash 2 Introduces Block-Diffusion Speculative Decoding to Speed Up GLM-5.3-Flash — songhan_mit · 2026-08-28
1 near-duplicate retellings: Lianhuiq