Block Diffusion (ICLR 2025 Oral) Marries Parallel Generation With KV Caching

alec_helbling · x · 2026-09-11

Block Diffusion (arXiv:2503.09573, ICLR 2025 Oral) interpolates between autoregressive and diffusion LMs. Diffusion LMs generate tokens in parallel but iterative unmasking updates token states, killing KV-cache reuse. Block Diffusion decodes blocks left-to-right while generating each block in parallel, restoring efficient caching and enabling flexible-length generation. The paper adds a recipe: efficient training, gradient-variance estimators, and data-driven noise schedules. It sets a new SOTA among diffusion models on LM benchmarks, with code and weights open-sourced.

Related event: Block Diffusion Brings KV-Cache Reuse to Diffusion LMs(2 posts)→

Original post →

More from Research

Research channel →