Continuous diffusion language models are making a comeback, Sander Dieleman writes

joao_gante · x · 2026-08-24

Sander Dieleman published a 44-minute read, "Continuous diffusion language models." He notes that language modeling has long been dominated by autoregression: token-by-token generation, teacher forcing for efficient parallel training, and extreme scalability — the recipe behind today's LLMs. But diffusion offers an alternative iterative generation path, inspired by early successes in audio and image domains.

He traces the history: early continuous-diffusion attempts for language were largely supplanted by fully discrete diffusion methods and went dormant, yet this year has seen a flurry of new research, suggesting a genuine revival. The post covers the historical perspective and technical aspects behind the comeback, presented as a subjective account open to dissenting views.

Original post →

More from Models

Models channel →