Qwen3.8-27B open-sourced with SGLang Day-0 support, 206 tok/s on RTX 5090
ying11231 · x · 2026-08-14
Alibaba's Qwen3.8-27B is open-sourced with SGLang Day-0 support. Achieves 206.1 tok/s decode on RTX 5090 (NVFP4+DSpark), 38.28 tok/s on DGX Spark. Native multimodal dense model, 27B params, outperforms Qwen3.7-Plus, 262K native context extendable to 1M. Apache 2.0.
Related event: Qwen3.8-27B Released, Tops Hugging Face Trends(20 posts)→
More from Models
- Qwen3.8-27B-FP8 on GH200: 10 concurrent streaming requests, first token in 10ms — MaziyarPanahi · 2026-08-15
- Alibaba Releases Qwen3.8-27B: 27B Parameters, Opus 4.6-Class Performance, Runs Locally — haider1 · 2026-08-15
- Qwen3.8-27B Open Source: 206 tok/s on RTX 5090 with SGLang Day-0 Support — cedric_chee · 2026-08-15
- Qwen Official: Qwen3.8-27B Significant Jump, Try It Out — Alibaba_Qwen · 2026-08-15
- Opus 4.6 Max Now Runs Locally, Enabling Cutting-Edge Research Anywhere — rand_longevity · 2026-08-15
- Hugging Face report: small models dominate real-world usage, Qwen leads local inference — huggingface · 2026-08-15