Custom fused sampler kernel adds 15% tok/s for DiffusionGemma on vLLM
imjustnewatai · x · 2026-09-23
Developer mmastrac reports a custom fused sampler kernel for vLLM delivering an extra 15% tok/s on DiffusionGemma for regular text generation — not the DiffusionGemma-as-Jev path — showing sampler kernel fusion remains untapped headroom in the inference stack.
Related event: Custom fused sampler kernel boosts vLLM throughput by 15%(2 posts)→
More from Infra
- Flash-dLLM accelerates diffusion LLMs up to 11x with IO-aware KV caching — MBZUAI · 2026-09-23
- vLLM v0.30.0 ships with 762 commits: watermarking, HiSparse, Model Runner V2 — vllm_project · 2026-09-23
- Sea first ASEAN company to adopt NVIDIA Vera Rubin as Nemotron spreads across Southeast Asia — NVIDIA Blog · 2026-09-23
- liuliu warns: claimed 4x-10x speedups over MLX or llama.cpp on Apple hardware are noise — teortaxesTex · 2026-09-23
- DeepSeek DSec cluster BoM estimated at ≤$20M; $1B could buy 50 clusters and 19M concurrent sandboxes — teortaxesTex · 2026-09-23
- B200 Rental Prices Hit All-Time High at $7.88 per GPU-Hour — sudoraohacker · 2026-09-23