Hy4 preview repo goes live: 256 experts per MoE layer, native MTP head, Gated DSA attention

aigclink · x · 2026-08-28

Tencent's Hy4 preview GitHub repo is now public. Architecture details: 770B total / 49B active parameters; a 78-layer backbone where the first layer uses dense FFN and the remaining 77 use MoE — each with 256 routed experts + 1 shared expert, activating top-8 per token. A native MTP layer (10B total, 0.7B active) is built in for speculative decoding. Inspired by DeepSeek and GLM, attention uses Gated DeepSeek Sparse Attention (Gated DSA) with IndexCache for cross-layer sparsity.

The repo includes vLLM/SGLang deployment, finetuning and quantization docs, with bilingual READMEs.

Related event: Tencent Open-Sources Hunyuan Hy4 Preview: 770B MoE with 1M Context(10 posts)→

Original post →

More from Models

Models channel →