Tencent's Sherry quantization shrinks Hy4 model size by 7x

QuixiAI · x · 2026-09-01

Tencent's Sherry quantization method compresses model weights to 1.25 bits. Applied to Hy4 Preview, it reduces size from 1.5TB to 214GB with minimal accuracy loss (e.g., MCP Atlas 83.7→83.2). It also enables stitching GPUs across machines.

Related event: Tencent's Sherry Quantization Shrinks Hy4 Model Sevenfold to 214GB(4 posts)→

Original post →

More from Research

Research channel →