Custom fused sampler kernel adds 15% tok/s throughput in vLLM for DiffusionGemma

generativist · x · 2026-09-23

Developer mmastrac implemented a custom fused sampler kernel for DiffusionGemma on vLLM, gaining an extra 15% tok/s in regular text inference (not in DiffusionGemma-as-Jev mode). Another concrete win for the inference serving stack.

Related event: Custom fused sampler kernel boosts vLLM throughput by 15%(2 posts)→

Original post →

More from Infra

Infra channel →