Hy3 Launches 1bit/4bit Quantized Versions
腾讯混元 · wechat · 2026-07-14
Tencent's Hunyuan team has performed quantization and deployment optimizations for the open-source Hy3 model, focusing on making the 295B flagship model runnable on more limited hardware.
Key highlights include:
- Releasing 1bit and 4bit GGUF quantized versions for use with the llama.cpp ecosystem.
- The 1bit version compresses the 598GB BF16 weights down to 85.5GiB, which the author says can be deployed on a single 96GB inference card.
- The 4bit version is approximately 169.9GiB, suitable for dual inference cards.
- Additionally, a GPTQ Int4 version is provided, which can be directly deployed using vLLM.
- The team also added support for Hy3's MTP speculative decoding to llama.cpp; when enabled, the acceptance rate is around 60%, delivering a speedup of roughly 50% for 1bit and nearly 60% for 4bit.
- Official Hugging Face links for the base model, quantized models, and patches are also provided.
Related event: Tencent Hunyuan ships 1-bit/4-bit quantized Hy3 for single-GPU deployment(5 posts)→
More from Infra
- Tabul AI launches Metal TreeSHAP to speed up Shapley values on Apple silicon — Scobleizer · 2026-07-22
- DeepSeek-V4-Flash tops out at 770 tok/s on one B300 in a vLLM batch test — Moreh · 2026-07-22
- NVIDIA starts shipping 102.4 Tbps Spectrum-6 switches for Vera Rubin AI factories — nvidia · 2026-07-22
- Apple publishes SOC 3 audit reports for Private Cloud Compute — throwfaraway4 · 2026-07-22
- Reddit GPU renters say existing platforms only give you two of three: code, recovery, fair billing — legendpizzasenpai · 2026-07-22
- The Sandboxing Manifesto: Secure Execution Environments for Agents — spirosoik · 2026-07-22