Google's DiffusionGemma: Retrofitting LLMs into Diffusion Models at 10% of the Cost
The Decoder · rss · 2026-08-09
Google DeepMind introduced DiffusionGemma, demonstrating that text diffusion models can be built without training from scratch. By retrofitting the existing Gemma 4 model, researchers achieved this using less than 10% of the original training budget.
Unlike traditional autoregressive approaches that predict tokens sequentially, this diffusion-based model generates 256 tokens in parallel, reaching speeds of approximately 1,500 tokens per second. However, in benchmark evaluations, its overall generation quality—particularly in reasoning tasks—still trails behind the original autoregressive model.
More from Models
- Meta Paper: Quantization Causes Overthinking in Reasoning Models, Fixable with Word Penalties — rohanpaul_ai · 2026-08-09
- NVIDIA Releases Nemotron-Parse-2.0 for Advanced Document Parsing — nvidia · 2026-08-09
- Leak: OpenAI's GPT-6 'Astra' to launch this month with 10T pre-train, vastly better than Fable — iruletheworldmo · 2026-08-09
- Alleged New GPT Image Model 'mona-lisa-1' Surfaces on Chatbot Arena — mark_k · 2026-08-09
- Gemini AI users request support for rendering code blocks in prompts — Novel-Nature-7741 · 2026-08-09
- Anthropic's Haiku Stagnates for a Year as OpenAI Accelerates Small Model Strategy — kimmonismus · 2026-08-09