Custom fused sampler kernel adds 15% tok/s for DiffusionGemma on vLLM

imjustnewatai · x · 2026-09-23

Developer mmastrac reports a custom fused sampler kernel for vLLM delivering an extra 15% tok/s on DiffusionGemma for regular text generation — not the DiffusionGemma-as-Jev path — showing sampler kernel fusion remains untapped headroom in the inference stack.

Related event: Custom fused sampler kernel boosts vLLM throughput by 15%(2 posts)→

Original post →

More from Infra

Infra channel →