Hy4 preview repo goes live: 256 experts per MoE layer, native MTP head, Gated DSA attention
aigclink · x · 2026-08-28
Tencent's Hy4 preview GitHub repo is now public. Architecture details: 770B total / 49B active parameters; a 78-layer backbone where the first layer uses dense FFN and the remaining 77 use MoE — each with 256 routed experts + 1 shared expert, activating top-8 per token. A native MTP layer (10B total, 0.7B active) is built in for speculative decoding. Inspired by DeepSeek and GLM, attention uses Gated DeepSeek Sparse Attention (Gated DSA) with IndexCache for cross-layer sparsity.
The repo includes vLLM/SGLang deployment, finetuning and quantization docs, with bilingual READMEs.
Related event: Tencent Open-Sources Hunyuan Hy4 Preview: 770B MoE with 1M Context(10 posts)→
More from Models
- GLM 5.3 Flash surges to 4th place in daily OpenRouter usage — Hesamation · 2026-08-28
- Tencent opens Hy4 preview weights: 770B MoE, 49B active, 1M context — Snoo26837 · 2026-08-28
- Similar benchmarks, double the size: Qwen3.8-Flash-Next needs 360GB vs DeepSeek-V4-Flash's 162GB lossless — vini542reddit · 2026-08-28
- Grok 4.6 hits 95% on GPQA Diamond, tied #1 and beating Opus 5 and GPT-5.6 — XFreeze · 2026-08-28
- Hands-on with Tencent Hy4 preview across 8 projects: better frontend taste, stable long-horizon tasks — vista8 · 2026-08-28
- PlayWorld Reveals Quality Gap in High-Scoring World Models — jiqizhixin · 2026-08-28