Independent Sweep Puts DiffusionGemma Peak Throughput at 3k tok/s, 3x Paper's Claim
bodonoghue85 · x · 2026-09-06
A researcher re-ran throughput benchmarks for DiffusionGemma after noticing the paper omitted batch size details. Using 1xH200, PG19 data, 17.56 toks/forward, 1k input / 8k output with vLLM, he measured peak throughput of 3k tok/s at fp8 and 2.7k at bf16 — roughly 3x higher than reported in the uno paper. The paper's claimed 1k tok/s peak is not accurate and needs updating. He also notes it's unusual to claim inference wins at bf16 rather than fp8, since no one serves at bf16 at scale, and fp8 delivered further gains as expected.
More from Infra
- Lightpanda: Open-Source Headless Browser for AI Agents, 11x Faster Than Chrome — Shruti_0810 · 2026-09-06
- Baseten's Philip Kiely Launches Inference Engineering Book, Plus Learning Resources — kmeanskaran · 2026-09-06
- Polygres turns your Postgres into a hybrid search context layer for AI agents — Scobleizer · 2026-09-06
- TCS may invest up to $7.4 billion with TPG in a gigawatt AI campus in Hyderabad — emmanuelvivier · 2026-09-06
- KV cache pressure tool shows vLLM's advertised 2M-token cache can retain 3M after fixes — t4a8945 · 2026-09-06
- Dev Ships Open-Source Sliding-Window Attention for HF LLMs, Hits 3.5MB KV Cache at 32K Context — ahsaor8 · 2026-09-06