Tencent compresses 1.5TB model to 200GB with minimal accuracy loss

ChrisGPT · x · 2026-08-29

Tencent has compressed the Hy4-preview model from 1.5TB to 200GB in GGUF format with negligible accuracy loss. The MIX-STQ10 method dynamically decides bit-width for each layer using calibration data: some layers are aggressively compressed to 1.31-bit, while sensitive layers remain near 2-bit.

This non-uniform quantization strategy significantly reduces error under the same compute budget, showcasing insane efficiency gains in low-bit inference.

Related event: Tencent compresses 1.5TB model to 200GB with near-lossless accuracy(2 posts)→

Original post →

More from Infra

Infra channel →