Tencent compresses 1.5TB model to 200GB with minimal accuracy loss
ChrisGPT · x · 2026-08-29
Tencent has compressed the Hy4-preview model from 1.5TB to 200GB in GGUF format with negligible accuracy loss. The MIX-STQ10 method dynamically decides bit-width for each layer using calibration data: some layers are aggressively compressed to 1.31-bit, while sensitive layers remain near 2-bit.
- MCP Atlas: 83.7 → 83.2
- SWE-Bench Multi: 82.9 → 81.3
- MRCR: 81.3 → 81.1
- IFBench: 73.5 → 72.5
This non-uniform quantization strategy significantly reduces error under the same compute budget, showcasing insane efficiency gains in low-bit inference.
Related event: Tencent compresses 1.5TB model to 200GB with near-lossless accuracy(2 posts)→
More from Infra
- Deploy Models with Kubernetes and TensorFlow Serving — Al_Grigor · 2026-08-29
- Serverless Deep Learning: Deploy Models on AWS Lambda — Al_Grigor · 2026-08-29
- Model Deployment: FastAPI, Docker, and Cloud Deployment — Al_Grigor · 2026-08-29
- Firefox & Chrome intend to ship support for JPEG XL — addyosmani · 2026-08-29
- Australia at 'sliding doors' moment to reap AI boom rewards via data centers — TobyWalsh · 2026-08-29
- Samsung details its Processing-in-Memory (PIM) architecture at Hot Chips 2026 — ingve · 2026-08-29