DFlash 2 speeds up GLM-5.3-Flash by 2.8x via block-diffusion speculative decoding
Lianhuiq · x · 2026-08-28
Inco AI released the DFlash 2 draft model for GLM-5.3-Flash, achieving a 2.8x speedup using speculative decoding. DFlash 2 employs a block-diffusion mechanism to predict a whole block of tokens in one pass and track coherent paths with a lightweight selector. Benchmarks run on SGLang with NVIDIA GB300 GPUs show lossless decoding where greedy outputs match the target model exactly.
Related event: Inco AI Releases DFlash 2 Draft Model, Speeding Up GLM-5.3-Flash by 2.8x(3 posts)→
More from Infra
- Polymarket prices 68% chance of a statewide data center moratorium — Polymarket · 2026-08-29
- Voters flip on data center bans once projects cover grid, water costs — Polymarket · 2026-08-29
- AI Data Centers Turn to Fuel Cells to Bypass Multi-Year Grid Delays — tengyanAI · 2026-08-29
- Observation: Why Are So Many People Suddenly Owning DGX Stations? — andrew_n_carr · 2026-08-29
- Understanding KV, Prefix, Prompt, and Semantic Caching in LLMs — blaizedsouza · 2026-08-29
- Offloading only "hot" MoE experts to VRAM boosts llama.cpp throughput 50% — nbvehrfr · 2026-08-29