Qwen3.8-27B launches with 1M context on a single GPU, vLLM ready

Alibaba_Qwen · x · 2026-08-14

Alibaba's Qwen team releases Qwen3.8-27B, a 27B-parameter dense hybrid-attention model with linear attention on 48 of 64 layers, 262K native context extendable to 1M, multimodal support, and built-in MTP draft head. It runs on a single Blackwell GPU (24.6 GiB in NVFP4) and is available on vLLM 0.17.0 from day one.

Related event: Qwen3.8-27B Released Open-Source, Tops Hugging Face Trending(36 posts)→

Original post →

More from Infra

Infra channel →