Tencent Hunyuan 770B compressed from ~1.5TB to ~214GiB with mixed quantization

Aiden_Tech_Ai · x · 2026-09-03

A leak about Tencent Hunyuan's 770B "Hy4 preview" claims mixed GGUF quantization shrinks it from 1.5TB to 214GiB with evals close to BF16. Caveats: it's per-layer mixed quantization, not a true "1-bit" model, and won't run natively on every laptop. Highlights include 49B active params per token, native 1M context, Apache 2.0 license, and vLLM/SGLang support. Big claims that now need independent testing.

Original post →

More from Infra

Infra channel →