Tencent's Hy4 Preview: 1.25-bit Quantization Cuts Model Size by 7x
ccerrato147 · x · 2026-09-01
Tencent's Hy4 model preview utilizes the Sherry quantization method, reducing weight precision to 1.25 bits. This shrinks the model size from 1.5TB to 214GB with negligible quality loss. It also enables stitching GPUs across multiple machines to work as one. Both GGUF quantized and original models are available.
Related event: Tencent's Sherry Quantization Shrinks Hy4 Model Sevenfold to 214GB(4 posts)→
More from Infra
- AWS Launches AgentCore, Accepting Multi-Cloud Reality — DavidLinthicum · 2026-09-01
- Shopify's Continual Learning Loop Cuts Serving Costs by 96% — yenkel · 2026-09-01
- Shopify Engineer Details Continual Learning for GraphQL Agent — Drewch · 2026-09-01
- Inference provider Wafer raises $40m at $200m+ valuation, rejecting offers — steph_palazzolo · 2026-09-01
- x402 Standard: Building a Native Payment Layer for the Agentic Economy — kleffew94 · 2026-09-01
- Custom HBM drives 5x LLM inference speed boost — BenBajarin · 2026-09-01