How Tencent Hunyuan packed a 770B model into 214 GiB with 5-bit-per-4-weights quantization
TencentHunyuan · x · 2026-09-22
Tencent Hunyuan's quantization team explains how Hy4 preview's 770B-parameter weights shrank from 1.5TB to 214 GiB while preserving capability and inference speed. Key details: the Sherry quantization algorithm, STQ10 storage format, and MIX-STQ10 mixed-precision allocation. Groups of four weights take values from {-d, 0, +d} with exactly one zero — effectively four weights in five bits, with zero positions used for further encoding. Parameter count is unchanged; only the weight representation changes.
More from Infra
- OpenAI removes Ultrafast tier from GPT-5.6 Sol in Codex, fueling GPT-6 Sol rumors — imjustnewatai · 2026-09-22
- Is upgrading from 2x to 4x RTX 3090 worth it for local LLM work? — fgoricha · 2026-09-22
- Qwen-Image local on a 24GB MacBook Pro takes 5-6 minutes per image — vista8 · 2026-09-22
- Cerebras CEO on Jensen Huang: a decade trading as 'nobody' before Nvidia made it — rohanpaul_ai · 2026-09-22
- 456GB DeepSeek v4.1 runs locally at 40 tok/s with Threadripper + dual RTX 6000 hybrid setup — HankYeomans · 2026-09-22
- Agent Substrate roadmap: sub-second suspend/resume runtime for dense agent deployments — rakyll · 2026-09-22