Tencent Hunyuan Hy3 Officially Released with Day-One Native vLLM Support
vllm_project · x · 2026-07-06
Tencent Hunyuan has released the official version of Hy3 (succeeding the Hy3 Preview). It natively supports the vLLM inference framework from day one and is equipped with tool call parsing, reasoning parsing, and MTP speculative decoding, verified on both NVIDIA and AMD hardware.
Hy3 is a MoE model designed for agentic workflows, coding, and long-horizon reasoning, released under the Apache 2.0 license. It features 295B total parameters with 21B activated, 192 experts with top-8 routing, GQA attention, and a 256K context window. It also includes a 3.8B parameter MTP speculative decoding layer, with both BF16 and FP8 versions available.
Related event: Tencent Open-Sources Hunyuan Hy3 for Agentic Workloads(26 posts)→
More from Infra
- PlanetScale launches sharded Postgres: 768 servers acting as one, 1PB scale — dhruv2038 · 2026-09-11
- Can a 7900 XTX 24GB run Qwen locally? Reddit seeks ROCm tok/s benchmarks — thenomadexplorerlife · 2026-09-11
- RTK Terminal Compression Cuts Tokens but Leaves Your AI Coding Bill Unchanged — Bartaseth · 2026-09-11
- SF Compute founder: buying compute is 'an absolutely awful experience' right now — IgorCarron · 2026-09-11
- SmolVM open-sources persistent computer infrastructure for agents that outlive chat sessions — aniketmaurya · 2026-09-11
- PyTorch Day Korea 2026 launches first offline conf, CFP closes Sept 13 — PyTorch · 2026-09-11