Alibaba's Qwen3.8 with 51B N-gram embeddings now available on SGLang

Alibaba_Qwen · x · 2026-08-26

Alibaba announced that Qwen3.8-Flash-Next is ready for deployment via SGLang. The architecture features a 125B main model paired with 51B N-gram embeddings, activating only 6B parameters per token.

Key highlights:

The model was trained using the Muon optimizer.

Related event: Alibaba Open-Sources Qwen3.8-Flash-Next: 125B Ultra-Sparse MoE Previewing Qwen4 Architecture(11 posts)→

Original post →

More from Infra

Infra channel →