DiffusionGemma hits 97 tok/s single-stream and 317 tok/s at concurrency 32; vLLM patches landing
GlennCameronjr · x · 2026-09-23
Developer mmastrac shared new DiffusionGemma inference benchmarks: 97 tok/s at concurrency 1 and 317 tok/s at concurrency 32 — running on a single Spark device. The author argues the diffusion-based approach has been overlooked for too long, and notes vLLM patches will land over the next few days, potentially making diffusion decoding a practical option for on-device inference.
More from Infra
- The Datacenter as a Computer: Revisiting Luiz André Barroso's Google Legacy — JeffDean · 2026-09-24
- AMD Shares Behind-the-Scenes Look at Building and Testing Helios Racks — wkmyrhang · 2026-09-24
- NVIDIA gifts Unsloth a DGX Station; team teases faster quantization and efficient RL support — danielhanchen · 2026-09-24
- NVIDIA Nemotron 3 Diarization on Baseten: 500+ real-time streams per RTX PRO 6000 — baseten · 2026-09-24
- Unsloth passes 500M downloads as local AI adoption outpaces its own predictions — danielhanchen · 2026-09-23
- MLSys 2027 opens call for papers with Song Han as PC Co-Chair — songhan_mit · 2026-09-23