Qwen3.8-27B Open Source: 206 tok/s on RTX 5090 with SGLang Day-0 Support
cedric_chee · x · 2026-08-15
Alibaba's Qwen3.8-27B is now open source, with Day-0 support from SGLang. On a single RTX 5090, it achieves 206.1 tok/s decode speed using NVFP4 and DSpark, and 38.28 tok/s on DGX Spark. The model excels in agentic planning and long-horizon tasks, earning the title 'king of small models'.
Related event: Qwen3.8-27B Released Open-Source, Tops Hugging Face Trending(36 posts)→
More from Models
- Qwen2.5-72B Leads Among Similar-Sized Models — solyarisoftware · 2026-08-15
- Harvey and Applied Compute train legal-specific model with SOTA accuracy at fraction of cost — rhythmrg · 2026-08-15
- Pestle 27B Ternary: first medical ternary model, Apache 2.0 — Individual-Dot5488 · 2026-08-15
- Frontier model prices drop across the board, open and closed source, making agents cheaper — markjeffrey · 2026-08-15
- RedNote AI lab's TEMPO RL method: 16B MoE scores >30% on ARC-AGI 3 — GregKamradt · 2026-08-15
- ChatGPT randomly inserts Delta while discussing Newton's laws — PtrPomorski · 2026-08-15