Tencent Hy4-preview shrunk to 214GB via Sherry quantization

alejandroll10 · x · 2026-09-02

Tencent announced shrinking Hy4-preview from 1.5TB to 214GB using the Sherry quantization method, packing weights to 1.25 bits each. It enables stitching existing GPUs across machines. Using MIX-STQ10, calibration data selects bit-width per layer (some down to 1.31-bit). Accuracy drop is minimal vs BF16: MCP Atlas 83.7→83.2, SWE-Bench 82.9→81.3. Weights and GGUFs are released.

Related event: Tencent's Sherry Quantization Shrinks Hunyuan Hy4 Preview from 1.5TB to 214GB(5 posts)→

Original post →

More from Models

Models channel →