Flash-dLLM accelerates diffusion LLMs up to 11x with IO-aware KV caching
MBZUAI · hf · 2026-09-23
MBZUAI introduces Flash-dLLM, a training-free inference acceleration framework for diffusion LLMs. It identifies GPU memory I/O as the dominant bottleneck in KV-cache-enabled dLLM inference and addresses it with an IO-aware fused KV-cache kernel, then adds a KV-cache-driven draft-and-verify decoding strategy where the dLLM serves as both drafter and verifier without an auxiliary model. On math reasoning and code generation benchmarks it delivers 5.1x and 11.0x speedups over the prior strongest baseline Elastic-Cache on GSM8K and HumanEval, while preserving quality and scaling to longer sequences and larger batches.
More from Infra
- Anthropic in early talks to lease up to 1GW from Apollo-backed data center developer — rohanpaul_ai · 2026-09-23
- Anthropic in early talks to lease up to 1GW of compute from Apollo-owned data center operator — rohanpaul_ai · 2026-09-23
- Inference startup: training wins inference workloads, models should improve with use — ypatil125 · 2026-09-23
- Apple M6 ANE memory bandwidth tops 150GB/s, matching GPU/CPU — AIFlow_ML · 2026-09-23
- Report: Anthropic in talks to cement control over more data centers — pstAsiatech · 2026-09-23
- Back-of-envelope math: MiMo's 1.27M RL rollouts could run in ~10 hours — teortaxesTex · 2026-09-23