Block Diffusion Brings KV-Cache Reuse to Diffusion LMs
Block Diffusion, an ICLR 2025 Oral paper, interpolates between autoregressive and diffusion language models by decoding left-to-right block by block with parallel generation within blocks, enabling efficient KV-cache reuse for diffusion LMs.
2026-09-11 ~ 2026-09-11 · 2 related posts
- Block Diffusion restores KV-cache efficiency for parallel-decoding diffusion LMs — alec_helbling · 2026-09-11
- Block Diffusion (ICLR 2025 Oral) Marries Parallel Generation With KV Caching — alec_helbling · 2026-09-11