Tencent Hunyuan Hy4-preview runs in vLLM day 0: 770B MoE with 1M context

TencentHunyuan · x · 2026-08-30

vLLM officially announced that Tencent Hunyuan's Hy4-preview is supported from day 0, verified on NVIDIA GPUs.

Architecture highlights:

Deployment: the FP8 build runs on 16×B200 or 8×B300; enabling VLLMENABLEHPCOPS=1 activates Tencent's HPC-Ops attention and MoE kernels (in vLLM main since Hy3), combined with MTP speculative decoding and hyv4 tool/reasoning parsers.

Original post →

More from Infra

Infra channel →