Tencent's Sherry quantization shrinks Hy4 model size by 7x
QuixiAI · x · 2026-09-01
Tencent's Sherry quantization method compresses model weights to 1.25 bits. Applied to Hy4 Preview, it reduces size from 1.5TB to 214GB with minimal accuracy loss (e.g., MCP Atlas 83.7→83.2). It also enables stitching GPUs across machines.
Related event: Tencent's Sherry Quantization Shrinks Hy4 Model Sevenfold to 214GB(4 posts)→
More from Research
- Alibaba Paper Proposes SkillZip Pro for Agent Compression — dair_ai · 2026-09-01
- Combining RMSE and R-squared for Better Forecast Model Evaluation — mdancho84 · 2026-09-01
- A data scientist's two R-squared mistakes that hurt his regression models for 2 years — mdancho84 · 2026-09-01
- Hardcore eval: Fixing gpt-oss harness and testing 320k cases — skeole · 2026-09-01
- Multilingual 3.7B MoE trained from scratch on a consumer GPU — Significant_Focus134 · 2026-09-01
- New PACT Benchmark Reveals Enterprise AI Compliance Failures Under Pressure — baseten · 2026-09-01