Tencent Hunyuan releases lightweight model with 85% weight compression

智东西 · wechat · 2026-09-01

Tencent released a lightweight version of its flagship Hunyuan Hy4preview model, compressing weights from 1.5TB to 214GB using proprietary Sherry sparse ternary quantization and MIX-STQ1.0 mixed precision strategies. Tests show the compressed model performs on par with the original in long-context understanding and retrieval, with a slight drop in math. Using the distributed inference framework prima.cpp, the team demonstrated running the model across heterogeneous devices (e.g., a 4090 laptop and A4000 server), achieving 6x faster inference than single-machine offload.

Related event: Tencent's Sherry Quantization Shrinks Hunyuan Hy4 Preview from 1.5TB to 214GB(5 posts)→

Original post →

More from Infra

Infra channel →