FULL STORY
Tencent Hunyuan Hy3: From Open Source to Quantization
Tencent open-sourced the Hunyuan Hy3 model, followed by local deployments, community testing, and the release of quantized versions, leading to a 68x spike in calls within a week.
2026-07-06 ~ 2026-07-16 · 6 episodes · 46 posts
Episode 1 · Tencent Open-Sources Hunyuan Hy3 for Agentic Workloads (2026-07-06, 26 posts)
Multiple posts say Tencent has released and open-sourced the formal version of Hunyuan Hy3, a 295B-parameter MoE model with 21B active parameters and 256K context. It stands out because it targets agentic workflows, tool use, coding, and long-context reasoning while keeping active compute relatively low, and it quickly landed in ecosystem products and inference stacks.
Key details
According to testingcatalog and other posters, Hy3 is the formal release that follows an April-end Hy3 Preview. Posts describe it as a 295B MoE model with 21B active parameters; testingcatalog also mentions a 3.8B MTP layer. Multiple posts add that it is released under Apache 2.0 and can be used commercially.
AdinaYakup says the model supports switchable reasoning modes, including no-thinking, low-intensity, and high-intensity settings, and that FP8 variants are available. aigclink further relays claimed reliability improvements over the previous generation, including hallucination rate dropping from 12.5% to 5.4%, commonsense error rate from 25.4% to 12.7%, and multi-turn failure rate from 17.4% to 7.9%. Because these figures are relayed by third-party posts in this cluster, they should be treated as reported claims here.
Several posts also highlight Tencent's own comparative positioning. iamfakhrealam says Tencent claims Hy3 can compete with open flagship models 2 to 5 times its size, and that in an internal blind test involving 270 experts on real work tasks it scored above GLM-5.1, with the largest gains in tool use, multi-step tasks, and web search.
Availability and ecosystem
vLLM says it natively supported Hy3 on day one, including tool-call parsing, reasoning parsing, and MTP speculative decoding, with validation on both NVIDIA and AMD hardware. Nous Research says Hy3 is available for free on Nous Portal for two weeks. Other posters add that the model is also accessible on OpenRouter, with iamfakhrealam reporting pricing at RMB 1 per million input tokens, RMB 4 per million output tokens, and RMB 0.25 for cache hits.
Reactions and early impressions
Discussion around Hy3 focused not only on benchmark scores but also on efficiency. kimmonismus notes that the model reportedly reaches its efficiency using standard GQA rather than sparse attention or MLA, arguing that this suggests further architectural optimization room still exists. Xianbao_QIAN says that in personal testing the model delivered strong opencode performance and could be especially competitive for deployments on 8x H200 servers.
rohanpaul_ai separately cites an atomic.chat local physics-simulation test in which Hy3 achieved physics quality close to Gemini 3.5 at roughly 35x lower cost. Sam Witteveen frames Hy3 as a serious new competitor to GLM 5.2, especially because it appears strong on many non-coding tasks while using about half the active parameters.
- 腾讯混元Hy3开源:295B总参数21B激活,Apache 2.0协议 — eliebakouch · 2026-07-06
- Hy3盲测胜出GLM-5.1,前端与CI/CD领域优势显著 — eliebakouch · 2026-07-06
- 腾讯Hy3模型实测:活跃参数约GLM 5.2一半,8卡H200可高效运行 — Xianbao_QIAN · 2026-07-06
- ModelScope发布Hy3:295B总参MoE模型,支持256K上下文 — bdsqlsz · 2026-07-06
- Tencent Hunyuan v3 (Hy3) MoE Hits HF Trending — tencent · 2026-07-06
- 腾讯开源混元Hy3:295B参数MoE,幻觉率降至5.4% — aigclink · 2026-07-06
- 腾讯开源混元HY3:295B MoE,支持可切换推理 — AdinaYakup · 2026-07-06
- 腾讯混元Hy3发布,侧重真实工作流场景评测 — airesearch12 · 2026-07-06
- 腾讯开源Hy3:295B MoE模型,推理编程媲美更大规模旗舰 — yshan2u · 2026-07-06
- 腾讯混元 Hy3 正式版开源,主打编程与智能体 — 智东西 · 2026-07-06
- 腾讯发布混元3:295B MoE盲测超GLM-5.1 — kimmonismus · 2026-07-06
- 腾讯发布 Hy3 正式版开源大模型 — testingcatalog · 2026-07-06
- 腾讯混元Hy3正式上线:295B MoE模型,主打编程与智能体 — AGI Hunt · 2026-07-06
- 腾讯混元Hy3正式版发布,vLLM首日原生支持 — vllm_project · 2026-07-06
- 腾讯混元Hy3本地物理模拟接近Gemini3.5水平 — rohanpaul_ai · 2026-07-07
- 腾讯混元 Hy3 295B MoE 模型在 Nous Portal 两周免费开放 — NousResearch · 2026-07-07
- Nous推出Portal订阅:一站式构建Hermes智能体 — NousResearch · 2026-07-07
- Tencent Releases Hy3 Model — bdsqlsz · 2026-07-07
- Tencent Hunyuan Hy3 Officially Released with Upgraded Yuanbao Agent Capabilities — 量子位 · 2026-07-07
- Tencent Hy3 Explained: A New Rival to GLM 5.2 — Sam Witteveen · 2026-07-07
Episode 2 · Tencent Hunyuan Hy3 Tested: Balancing Speed and Deep Reasoning (2026-07-07, 2 posts)
Developers tested Tencent Hunyuan Hy3, praising its 256K context, stable long conversations, and strong instruction following in coding and agent tasks. The model stands out by effectively combining fast responses with deep reasoning.
- Testing Tencent Hunyuan Hy3: Fast Responses and Solid Reasoning — HeyAmit_ · 2026-07-07
- Testing Tencent Hunyuan Hy3: Stable, 256K Context — FellMentKE · 2026-07-07
Episode 3 · Tencent Releases Open-Source Hunyuan Hy3 Model for Coding and Agents (2026-07-08, 9 posts)
Tencent officially released the open-source large language model Hunyuan Hy3. Recognized by many as one of the most notable open-source releases of the year, it boasts strong coding and agentic capabilities alongside high cost-effectiveness.
Core Specs and Capabilities
The model utilizes a 295B parameter Mixture-of-Experts (MoE) architecture, activating only 21B parameters per inference, and supports a maximum context window of 256K. According to posts, Hunyuan Hy3 is specifically built for agentic workflows and excels at agent-style coding tasks, capable of autonomously breaking down tasks, writing code, explaining logic, and iterating. Furthermore, the model demonstrates highly competitive performance across multiple coding, agent, and reasoning benchmarks.
Open Source and Trial Resources
Hunyuan Hy3 is released under the Apache 2.0 open-source license, with weights available on platforms like GitHub, Hugging Face, ModelScope, and GitCode. The official release includes a two-week free trial access via Hy AI Studio, and it is available on platforms like OpenRouter at prices significantly lower than competing models.
- Tencent Releases Open-Source Model Hy3 — socialwithaayan · 2026-07-08
- Tencent Releases Open-Source Hunyuan Hy3 — HeyZoyaKhan · 2026-07-09
- Tencent Hunyuan Hy3 Benchmarks and Trial Links — HeyZoyaKhan · 2026-07-09
- Hunyuan Hy3 Free Trial and Resources — HeyZoyaKhan · 2026-07-09
- Tencent Releases Open-Source MoE Model Hunyuan Hy3 — eyishazyer · 2026-07-09
- Tencent Releases Open-Source Model Hy3 — eyishazyer · 2026-07-09
- Hunyuan Hy3 Targets Coding and Agents — nikola_mr64990 · 2026-07-09
- Tencent Hy3 Released: Affordable and Great for Coding — kimmonismus · 2026-07-10
- Tencent Hy3 Release Info Update — kimmonismus · 2026-07-10
Episode 4 · Tencent HY3 Model Successfully Deployed on 128GB Mac (2026-07-11, 2 posts)
A developer successfully deployed Tencent's open-source HY3 model on a 128GB Mac. The test detailed the configuration process, quantization files, and overall performance, proving the feasibility of running massive models locally.
- Testing Tencent HY3 Local Deployment on a 128GB Mac — returnity · 2026-07-11
- Testing HY3 on a 128GB Mac — amerkthetrippyone · 2026-07-11
Episode 5 · Tencent Hunyuan ships 1-bit/4-bit quantized Hy3 for single-GPU deployment (2026-07-14, 5 posts)
Tencent Hunyuan shipped quantization and deployment optimizations for its open-source flagship Hy3, a 295B-parameter MoE, releasing 1-bit and 4-bit GGUF variants that the team says can run on a single GPU through inference solutions like llama.cpp. For those watching local LLM deployment, the core takeaway is bringing an otherwise high-barrier flagship model down to more accessible hardware.
Key details
Per Hunyuan's official post and several reposts, the release provides 1-bit and 4-bit GGUF quantized versions of Hy3, focused on inference usability. The team recommends pairing them with llama.cpp and mentions MTP support; with quantization and the matching inference setup, Hy3 can be deployed on a single GPU. Reposts also relay the claim that Hy3 leads among same-size models and competes with larger flagships.
Background and impact
Hy3 is positioned as a 295B flagship MoE. These posts do not include full benchmark results, specific single-GPU configurations, or measured performance figures, so the confirmable increment here is mostly the low-bit releases and the single-GPU deployment direction — its most practical selling point for real-world deployment.
- Hy3 Launches 1bit/4bit Quantized Versions — 腾讯混元 · 2026-07-14
- Tencent Hunyuan Releases Quantized Hy3 — victormustar · 2026-07-14
- Hy3 Launches 1-bit and 4-bit Versions — QuixiAI · 2026-07-14
- Tencent Hunyuan Launches Single-GPU Quantized Hy3 — huggingface · 2026-07-14
- Hunyuan Releases Low-Bit Version of Hy3 — _akhaliq · 2026-07-15
Episode 6 · Tencent Hunyuan Hy3 Usage Jumps 68x in One Week (2026-07-15, 2 posts)
Tencent said Hunyuan Hy3’s total API calls surged more than 68x versus the previous-generation Hy2 just one week after launch. In WorkBuddy’s manual model-selection scenarios, 60% of users chose Hy3, indicating strong early user adoption.
- Tencent Hunyuan Hy3 API Calls Surge 68x — 腾讯混元 · 2026-07-15
- Tencent Hunyuan Hy3 Sees Rapid Growth After Launch — Med1_Ai · 2026-07-16