Tencent Hunyuan Releases Quantized Hy3
victormustar · x · 2026-07-14
Tencent Hunyuan has released 1-bit and 4-bit versions of Hy3.
- This is a flagship MoE model with 295B parameters. Officially claimed to be leading in its size class, it is competitive with much larger flagship models.
- Quantized versions can run via llama.cpp. With MTP enabled, it can be deployed on a single GPU, emphasizing strong intelligence capabilities at lower hardware costs.
- The release is under the Apache 2.0 open-source license, making it suitable for commercial use, and includes a 2-week free API.
Related event: Tencent Hunyuan ships 1-bit/4-bit quantized Hy3 for single-GPU deployment(5 posts)→
More from Infra
- LLM Serving Metrics Thread: Why TPOT and Uptime Make or Break User Experience — abhijithneil · 2026-09-11
- PlanetScale launches sharded Postgres: 768 servers acting as one, 1PB scale — dhruv2038 · 2026-09-11
- Can a 7900 XTX 24GB run Qwen locally? Reddit seeks ROCm tok/s benchmarks — thenomadexplorerlife · 2026-09-11
- RTK Terminal Compression Cuts Tokens but Leaves Your AI Coding Bill Unchanged — Bartaseth · 2026-09-11
- SF Compute founder: buying compute is 'an absolutely awful experience' right now — IgorCarron · 2026-09-11
- SmolVM open-sources persistent computer infrastructure for agents that outlive chat sessions — aniketmaurya · 2026-09-11