Custom fused sampler kernel adds 15% tok/s throughput in vLLM for DiffusionGemma
generativist · x · 2026-09-23
Developer mmastrac implemented a custom fused sampler kernel for DiffusionGemma on vLLM, gaining an extra 15% tok/s in regular text inference (not in DiffusionGemma-as-Jev mode). Another concrete win for the inference serving stack.
Related event: Custom fused sampler kernel boosts vLLM throughput by 15%(2 posts)→
More from Infra
- Google Colab folds into AI plans: priority GPUs and background execution for Ultra users — Arindam_1729 · 2026-09-23
- AWS Open-Sources Strands Harness, Claims 28% Fewer Agent Tokens — shashib · 2026-09-23
- DIY multi-GPU cooling: case airflow tuning drops temps from 80C+ to 68C, no liquid cooling needed — HankYeomans · 2026-09-23
- Alibaba accelerates global AI push with new data centers across Europe and the Middle East — Polymarket · 2026-09-23
- You run kernels, not models: why the same model and GPU can perform wildly differently — Roger_M_Taylor · 2026-09-23
- Apple's Mac mini and Mac Studio get major AI-focused performance leap — BLUECOW009 · 2026-09-23