Tencent Hunyuan 770B compressed from ~1.5TB to ~214GiB with mixed quantization
Aiden_Tech_Ai · x · 2026-09-03
A leak about Tencent Hunyuan's 770B "Hy4 preview" claims mixed GGUF quantization shrinks it from 1.5TB to 214GiB with evals close to BF16. Caveats: it's per-layer mixed quantization, not a true "1-bit" model, and won't run natively on every laptop. Highlights include 49B active params per token, native 1M context, Apache 2.0 license, and vLLM/SGLang support. Big claims that now need independent testing.
More from Infra
- Buying a refurbished 8xA100 server to colocate and rent out: one user's GPU math — Alarming-Ad8154 · 2026-09-03
- AI Server PCB Drill Bits Emerge as Hidden Supply Chain Bottleneck — tengyanAI · 2026-09-03
- Hidden consumable under AI servers: PCB drill-bit demand is growing far faster than board volumes — tengyanAI · 2026-09-03
- Vera Rubin NVL72 hits up to 10X tokens per MW vs GB200 on DeepSeek R1, CoreWeave data shows — Beth_Kindig · 2026-09-03
- Cloudflare's Cache Transcoding shrinks cached text assets to ~1/3 with Zstandard — arpit_bhayani · 2026-09-03
- Microcenter shelf suddenly stocked with dozens of RTX 5090s — is the GPU shortage easing? — OvertaxedOne · 2026-09-03