Qwen3.8-27B launches with 1M context on a single GPU, vLLM ready
Alibaba_Qwen · x · 2026-08-14
Alibaba's Qwen team releases Qwen3.8-27B, a 27B-parameter dense hybrid-attention model with linear attention on 48 of 64 layers, 262K native context extendable to 1M, multimodal support, and built-in MTP draft head. It runs on a single Blackwell GPU (24.6 GiB in NVFP4) and is available on vLLM 0.17.0 from day one.
Related event: Qwen3.8-27B Released Open-Source, Tops Hugging Face Trending(36 posts)→
More from Infra
- Texas Tightens AI Data Center Approvals: Audits Required for Power, Water, and Community Impact — rohanpaul_ai · 2026-08-15
- Qwen3.8-27B crawls at 5 tokens/s on 8GB VRAM + 32GB RAM: best config? — SoAp9035 · 2026-08-15
- Harvey Trains Custom Model to Cut Costs and Boost Quality in Legal Review — ypatil125 · 2026-08-15
- TrendForce Raises AI Accelerator Shipment Forecast to 31% YoY Growth — Beth_Kindig · 2026-08-15
- NInfer Adds Day-0 Support for Qwen3.8-27B, Hits ~200 tok/s on RTX 5090 — FormOne2615 · 2026-08-15
- Orion-16B passes 100B tokens, largest LLM pretrained with decentralized compute — const_reborn · 2026-08-15