SGLang updates Qwen3.8-27B recipes, hitting 206 tok/s on RTX 5090

ying11231 · x · 2026-08-18

SGLang released updated recipes for running Qwen3.8-27B on RTX 5090 and RTX Pro 6000. The update adds variants for non-speculative decoding, MTP, and DSpark, alongside high throughput and low latency options. Benchmarks show 206.1 tok/s decode on a single RTX 5090 using NVFP4 plus DSpark. The recipes serve as starting points, requiring users to tune flags and configs for specific use cases.

Original post →

More from Infra

Infra channel →