Tencent's Hy4 Preview: 1.25-bit Quantization Cuts Model Size by 7x

ccerrato147 · x · 2026-09-01

Tencent's Hy4 model preview utilizes the Sherry quantization method, reducing weight precision to 1.25 bits. This shrinks the model size from 1.5TB to 214GB with negligible quality loss. It also enables stitching GPUs across multiple machines to work as one. Both GGUF quantized and original models are available.

Related event: Tencent's Sherry Quantization Shrinks Hy4 Model Sevenfold to 214GB(4 posts)→

Original post →

More from Infra

Infra channel →