SGLang updates Qwen3.8-27B recipes, hitting 206 tok/s on RTX 5090
ying11231 · x · 2026-08-18
SGLang released updated recipes for running Qwen3.8-27B on RTX 5090 and RTX Pro 6000. The update adds variants for non-speculative decoding, MTP, and DSpark, alongside high throughput and low latency options. Benchmarks show 206.1 tok/s decode on a single RTX 5090 using NVFP4 plus DSpark. The recipes serve as starting points, requiring users to tune flags and configs for specific use cases.
More from Infra
- Cursor releases Origin: a Git version optimized for AI agents — nicolascraske · 2026-08-18
- $10T in AI datacenter capex blocked: GPUs need 230GW, US grid delivers ~100GW — PeterDiamandis · 2026-08-18
- Power Grid Bottlenecks Stall $10T AI Data Center Buildout — PeterDiamandis · 2026-08-18
- Discussion: Running bots locally on Linux VMs — Daniel_Farinax · 2026-08-18
- PotatoMesh: Federated Dashboard for Visualizing LoRa Node Positions — tom_doerr · 2026-08-18
- llama.cpp adaptive MTP PR speeds up code generation by up to 100% — Look_0ver_There · 2026-08-18