vLLM Integrates Tencent Hunyuan Production Kernels, Boosting Hopper Inference

vllm_project · x · 2026-07-06

The vLLM project announced that serving the Hunyuan3 (Hy3) model on NVIDIA Hopper is now significantly optimized. The HPC-Ops attention and MoE backend, originally used by Tencent's Hunyuan AI Infra team in their production environment, has become a first-class citizen backend in the vLLM main branch.

The solution includes a step-wise load-balancing decode scheduler and a fully fused FP8 MoE pipeline. It achieves up to a 2.95x speedup on mixed-length decoding compared to static split-KV scheduling, and reduces Hy3's TTFT by about 24% and TPOT by roughly 17%. It can be used without forking or modifying the source code.

Original post →

More from Infra

Infra channel →