Block Diffusion restores KV-cache efficiency for parallel-decoding diffusion LMs
alec_helbling · x · 2026-09-11
- Diffusion LMs generate multiple tokens in parallel, but iterative unmasking repeatedly updates token states, limiting KV-cache reuse.
- Block Diffusion decodes blocks left-to-right while generating each block in parallel, restoring efficient caching and combining diffusion parallelism with AR-style cache efficiency.
Related event: Block Diffusion Brings KV-Cache Reuse to Diffusion LMs(2 posts)→
More from Research
- Tao and Fields Medalists' two objections to AI in math, and why they're weak — RexDouglass · 2026-09-12
- Conjectures launches Bittensor bounties paying TAO for cracking math problems open 30-80 years, judged by machine — markjeffrey · 2026-09-12
- FADA (CoRL 2026) open-sourced: humanoid robots adapt to new conditions from 2 minutes of experience — GuanyaShi · 2026-09-12
- Intel's silicon photonics couplers hit 1-1.5 dB IL, with visible epoxy delamination flaws — jwt0625 · 2026-09-12
- CPO paper criticized for vague DLW-to-PIC coupling description: 'such as TCB' — jwt0625 · 2026-09-12
- Fly connectome trained to play a Flappy Bird–style game — TinfoilTricorn · 2026-09-12