Tencent compresses 1.5TB model to 200GB with minimal benchmark loss via layer-wise calibration

ChrisGPT · x · 2026-08-29

Tencent demonstrated extreme model compression efficiency, reducing Hy4 preview from 1.5TB to 200GB with negligible impact on benchmarks. The core innovation involves using calibration data to dynamically determine compression intensity per layer: less sensitive layers are aggressively pushed down to 1.31 bit, while sensitive layers remain near 2 bit to prevent collapse. This layer-sensitivity-based differential quantization strategy yields massive efficiency gains.

Related event: Tencent compresses 1.5TB model to 200GB with near-lossless accuracy(2 posts)→

Original post →

More from Infra

Infra channel →