Diffusion will be everywhere: why text diffusion models may replace autoregressive LLM inference

akbirthko · x · 2026-10-02

Varun Neal's blog post "Diffusion will be everywhere" argues text diffusion is poised to spread across the LLM ecosystem:

Core claim: Diffusion offers advantages over autoregressive models for inference efficiency, test-time scaling, and RL — and converting strong AR checkpoints into diffusion models is now cheap enough for any well-resourced lab to adopt.

How it works: Generation starts from a canvas of random tokens; each denoising step predicts every position but only commits the most confident ones, so P tokens are produced in K forward passes (K≪P). DiffusionGemma, for example, generates a 256-token canvas in at most 48 steps, averaging 12 with adaptive stopping.

The tradeoff: Diffusion spends K times more compute but cuts memory reads to P/K. Since loading bytes from memory is hundreds of times slower than computing on GPUs, this trade favors diffusion when decoding is memory-bound.

The author predicts rapid ecosystem-wide adoption as the conversion process matures.

Original post →

More from Models

Models channel →