Helion-powered vLLM linear backend on Hopper yields 1.11-1.18x kernel speedups
lmoroney · x · 2026-10-03
A new PyTorch blog post from Red Hat, Meta, and the vLLM teams shows Helion powering a linear backend for vLLM on NVIDIA Hopper:
- One Helion GEMM implementation covers Standard, Split-K, and Swap-AB variants, with an ahead-of-time autotuner picking the best config per shape
- Hybrid dispatch: Helion handles small token counts (up to 32), default CUTLASS or DeepGEMM handles larger ones
- Results: geometric-mean kernel speedups of 1.11x-1.18x over default libraries, and >10% end-to-end throughput gains on some workloads
- Code and autotuning tooling ship in their vLLM fork, with an RFC open for upstream discussion
The post frames it as a clean ML systems teaching case: express the algorithm once at a high level, let autotuning own shape-specific work, and keep the hot path under CUDA Graphs so tuning wins survive into serving.
Related event: PyTorch Helion Kernels Boost vLLM Inference Throughput(2 posts)→
More from Infra
- INT21: 2 engineers direct AI to build 20 inference engines in 2 weeks, up to 2.4× faster than SGLang — bingxu_ · 2026-10-03
- Traversal's 5 Levels of Self-Driving Production: Why Coding Agents Make Ops Harder — AI Engineer · 2026-10-03
- How DatologyAI Generated 12 Trillion Synthetic Tokens — And Fixed 4 Pipeline Bottlenecks — AI Engineer · 2026-10-03
- Price-Insensitive Buyer With Billions Seeks 500MW-2GW of Powered Data Center Land — JohnnyNi13 · 2026-10-03
- Running 256k-context open models on 2x RTX 3090 for months: a home server LLM retrospective — knighty1981 · 2026-10-03
- Prime Intellect compresses MLA KV cache in NVFP4, fitting ~50% more tokens than FP8 — TheZachMueller · 2026-10-03