Qwen3.8-27B hits 206 tok/s on single RTX 5090 via SGLang
StefanoGogioso · x · 2026-08-17
SGLang announces Day-0 support for the Qwen3.8-27B model. Benchmarks show decoding speeds of 206.1 tok/s on a single RTX 5090 with NVFP4 and DSpark, and 38.28 tok/s on DGX Spark. The model excels at agentic planning and long-horizon tasks, with strong capabilities in 3D, game design, and coding.
More from Infra
- Bittensor Subnet 118 Adds Ultra-Cheap Inference, Joining Major AI Providers — markjeffrey · 2026-08-17
- Meta to rely on Nvidia Blackwell, AMD Helios in 2026, accelerate custom MTIA in 2027 — Beth_Kindig · 2026-08-17
- Stripe to Acquire OpenRouter for Over $7B, 5.4x May Valuation — rohanpaul_ai · 2026-08-17
- Wici One claims to solve local VRAM limits via NVMe offloading — Torodaddy · 2026-08-17
- antirez Optimizes DwarfStar: 170 t/s Generation and 22k tokens/s Prefill on Station — antirez · 2026-08-17
- CoreWeave Revenue Outpaces Big Cloud Early Stages Amid AI Pivot — a16z · 2026-08-17