Tencent Hunyuan releases lightweight model with 85% weight compression
智东西 · wechat · 2026-09-01
Tencent released a lightweight version of its flagship Hunyuan Hy4preview model, compressing weights from 1.5TB to 214GB using proprietary Sherry sparse ternary quantization and MIX-STQ1.0 mixed precision strategies. Tests show the compressed model performs on par with the original in long-context understanding and retrieval, with a slight drop in math. Using the distributed inference framework prima.cpp, the team demonstrated running the model across heterogeneous devices (e.g., a 4090 laptop and A4000 server), achieving 6x faster inference than single-machine offload.
More from Infra
- Sequoia invests $100M in Form Energy's iron-air batteries for the AI grid — DavidCahn6 · 2026-09-02
- DuckDB async I/O boosts speed: HF dataset reads 2-3x faster — vanstriendaniel · 2026-09-02
- GitHub Auto-Ban on Core Contributor Sparks Developer Concerns — braelyn_ai · 2026-09-02
- einopx: A unified JAX-native array pattern API — A_K_Nain · 2026-09-02
- Data centers provide nearly 50% of property tax in richest US county — aleximm · 2026-09-02
- Google details its full AI stack for developers — fhinkel · 2026-09-02