Tencent Hunyuan AngelSlim: Compressing Hy4 Model to 214GB with Heterogeneous Inference
腾讯混元 · wechat · 2026-09-01
Tencent's Hunyuan team released the AngelSlim quantization scheme, successfully compressing the Hy4-preview model (originally 1.5TB) down to 214GB, significantly lowering hardware barriers. Key technologies include:
- Sherry Sparse Ternary Quantization: Aggressively compresses MoE routing experts to an average of 1.25 bits per weight.
- MIX-STQ10 Mixed Precision Strategy: Dynamically allocates bits per layer based on sensitivity, balancing compression rate and accuracy.
Evaluations show minimal capability loss in long-context, math, and coding tasks, with some retrieval tasks even outperforming the original. Furthermore, the team validated heterogeneous device collaborative inference with prima.cpp, distributing the 214GB model across a laptop (RTX 4090) and a server (4x A4000). The setup achieved 1.02 token/s, proving that flagship models can run without expensive homogeneous clusters.
Relevant GGUF libraries and code have been open-sourced.
More from Infra
- 40-nm Memristor Chip Turns Conductance Drift Into a Feature, Beats A100 by 50-480x — maier_ak · 2026-09-01
- 40nm Neural-Dynamics Chip Uses Conductance Drift for 2.12ms Iteration Latency — maier_ak · 2026-09-01
- Qwen3.8 Flash hits 415 tok/s on dual DGX Sparks — NVIDIAAI · 2026-09-01
- OpenAI's 'Jalapeno' Chip Revealed: 1500 Tokens/s Throughput — firstadopter · 2026-09-01
- TensorSharp vs llama.cpp: Qwen 3.8 Flash Next Benchmarks — fuzhongkai · 2026-09-01
- Samsung shifts to 8-layer HBM4E for Nvidia with ~20% higher speed spec — 创业邦 · 2026-09-01