Tencent Hunyuan AngelSlim: Compressing Hy4 Model to 214GB with Heterogeneous Inference

腾讯混元 · wechat · 2026-09-01

Tencent's Hunyuan team released the AngelSlim quantization scheme, successfully compressing the Hy4-preview model (originally 1.5TB) down to 214GB, significantly lowering hardware barriers. Key technologies include:

Evaluations show minimal capability loss in long-context, math, and coding tasks, with some retrieval tasks even outperforming the original. Furthermore, the team validated heterogeneous device collaborative inference with prima.cpp, distributing the 214GB model across a laptop (RTX 4090) and a server (4x A4000). The setup achieved 1.02 token/s, proving that flagship models can run without expensive homogeneous clusters.

Relevant GGUF libraries and code have been open-sourced.

Original post →

More from Infra

Infra channel →