How Tencent Hunyuan packed a 770B model into 214 GiB with 5-bit-per-4-weights quantization

TencentHunyuan · x · 2026-09-22

Tencent Hunyuan's quantization team explains how Hy4 preview's 770B-parameter weights shrank from 1.5TB to 214 GiB while preserving capability and inference speed. Key details: the Sherry quantization algorithm, STQ10 storage format, and MIX-STQ10 mixed-precision allocation. Groups of four weights take values from {-d, 0, +d} with exactly one zero — effectively four weights in five bits, with zero positions used for further encoding. Parameter count is unchanged; only the weight representation changes.

Original post →

More from Infra

Infra channel →