Tencent Hunyuan Hy4 preview: 770B total/49B active, 1M context, Apache 2.0, day-0 vLLM
aftahi_ai · x · 2026-09-05
Tencent's Hunyuan Hy4 preview packs 770B total / 49B active params, native 1M context, Apache 2.0 licensing, day-0 vLLM and SGLang support, and official FP8 weights. Mixed per-layer quantization cuts memory from 1.5TB to 214GiB with reported accuracy barely moving vs BF16. A demo shows it turning a simple prompt into a playable game — a 770B-class model with a far lower deployment barrier.
More from Infra
- Burn Bar for Omarchy visualizes Claude/Codex token burn, quotas and GPU load locally — DanWahlin · 2026-09-05
- Scaling wall? Reddit argues test-time compute is the industry's new playbook — erdematar · 2026-09-05
- Qwen3.8 27B Quant Fits 24GB VRAM at 100k Context, Sparking Local Model Profit-Threat Debate — ChopSticksPlease · 2026-09-05
- Speechify CEO on self-built data centers, ElevenLabs leapfrog, and the $15M AI talent war — 20VC · 2026-09-05
- He Uses Local LLMs Like a 3D Printer: 12 Games, 29 Mods and Countless Tools Built Solo — Quebber · 2026-09-05
- Zhipu monetizes compute at $8-10M/MW, 5x below Anthropic and OpenAI's $40-50M/MW — zephyr_z9 · 2026-09-05