PyTorch Helion DSL integrated into vLLM beats CUTLASS and DeepGEMM with 10%+ throughput gains
PyTorch · x · 2026-10-03
PyTorch showcased Helion, its kernel DSL, by integrating it into vLLM's linear backend.
- A single Helion GEMM implementation covers multiple algorithmic variants — Standard GEMM, Split-K, and Swap-AB — with the best variant and config selected automatically per shape.
- On NVIDIA Hopper GPUs, per-shape tuning combined with hybrid dispatch outperforms vLLM's default CUTLASS and DeepGEMM backends across the evaluated models.
- The result is consistent end-to-end gains, with more than 10% throughput improvement on some workloads, in production inference settings.
- Written by Sean Chen (Red Hat) and Shangdi Yu (PyTorch, Meta).
More from Infra
- Epoch AI: chips shipped through 2027 could run up to hundreds of millions of concurrent frontier agents — KyeGomezB · 2026-10-03
- SK Hynix's Solidigm Reportedly Weighing $150 Billion IPO, Nearly Triple Arm's — Beth_Kindig · 2026-10-03
- ezyang's DeepSeek-V3 series part 3: roofline analysis for training DSv3 on Hopper — ezyang · 2026-10-03
- Developer slams DeepInfra's 33% input token price hike with just 2 days notice — julianharris · 2026-10-03
- Analyst: 5GW Live Compute Could Bring SpaceX ~$200B a Year, Supporting a $2T+ Valuation — JOBhakdi · 2026-10-03
- GPT-6.1 Sol reportedly under heavy load; capacity expansion to nearly double serving speed — NandaVegg · 2026-10-03