Qwen Team Analyzes Qwen3.8-Next Architecture Design

Qwen · hf · 2026-09-01

The Qwen team released a post on the design of the Qwen3.8-Next architecture. Qwen3.8-Flash-Next uses a sparse mixture-of-experts architecture, combining hybrid gated delta-net and sparse attention layers, gated residual branches, and off-accelerator n-gram embeddings to improve efficiency, capability, and training stability.

Original post →

More from Models

Models channel →