Google's DiffusionGemma: Retrofitting LLMs into Diffusion Models at 10% of the Cost

The Decoder · rss · 2026-08-09

Google DeepMind introduced DiffusionGemma, demonstrating that text diffusion models can be built without training from scratch. By retrofitting the existing Gemma 4 model, researchers achieved this using less than 10% of the original training budget.

Unlike traditional autoregressive approaches that predict tokens sequentially, this diffusion-based model generates 256 tokens in parallel, reaching speeds of approximately 1,500 tokens per second. However, in benchmark evaluations, its overall generation quality—particularly in reasoning tasks—still trails behind the original autoregressive model.

Original post →

More from Models

Models channel →