DiffusionGemma hits 97 tok/s single-stream and 317 tok/s at concurrency 32; vLLM patches landing

GlennCameronjr · x · 2026-09-23

Developer mmastrac shared new DiffusionGemma inference benchmarks: 97 tok/s at concurrency 1 and 317 tok/s at concurrency 32 — running on a single Spark device. The author argues the diffusion-based approach has been overlooked for too long, and notes vLLM patches will land over the next few days, potentially making diffusion decoding a practical option for on-device inference.

Original post →

More from Infra

Infra channel →