Inco AI Releases DFlash 2 Draft Model to Speed Up GLM-5.3-Flash Inference

Inco AI has released DFlash 2, a draft model that uses block-diffusion-based speculative decoding to accelerate zai-org/GLM-5.3-Flash inference by 2.8x.

2026-08-28 ~ 2026-08-28 · 2 related posts

1 near-duplicate retellings: Lianhuiq