Block Diffusion (ICLR 2025 Oral) Marries Parallel Generation With KV Caching
alec_helbling · x · 2026-09-11
Block Diffusion (arXiv:2503.09573, ICLR 2025 Oral) interpolates between autoregressive and diffusion LMs. Diffusion LMs generate tokens in parallel but iterative unmasking updates token states, killing KV-cache reuse. Block Diffusion decodes blocks left-to-right while generating each block in parallel, restoring efficient caching and enabling flexible-length generation. The paper adds a recipe: efficient training, gradient-variance estimators, and data-driven noise schedules. It sets a new SOTA among diffusion models on LM benchmarks, with code and weights open-sourced.
Related event: Block Diffusion Brings KV-Cache Reuse to Diffusion LMs(2 posts)→
More from Research
- Synthetic Morphology Suggests Non-Physicalist Models of Mind Can Be Empirically Tested — ZeroStateReflex · 2026-09-12
- Apple's Internalized Visual Thinking Drops the Paint-the-Future Pipeline for ~5x Faster Video Reasoning — jiqizhixin · 2026-09-12
- Skild AI founder explains why robotics data needs four sources, each flawed — deepakpathak · 2026-09-12
- Sony CSL's Frank Nielsen releases guaranteed arbitrary-precision approximations of Fisher-Rao geodesic distance — FrnkNlsn · 2026-09-12
- LittleLearner: a 5B model trained from scratch on a K-5-only corpus tests education data limits — repligate · 2026-09-12
- Conjectures launches Bittensor bounties paying TAO for cracking math problems open 30-80 years, judged by machine — markjeffrey · 2026-09-12