RTX 5000 Pro Runs Qwen3.8 at Unbeatable Value

Valuable-Run2129 · reddit · 2026-08-24

A user shared impressive benchmarks running Qwen3.8-27B on an RTX 5000 Pro (48GB VRAM): 130 t/s decode, 4000 t/s prefill, and 5 concurrent requests. Using SGLang, prefixes are offloaded to storage to prevent reprocessing in sub-agentic workflows. The author argues this model drastically increases local compute value, rivaling a Claude subscription and driving up the card's resale price.

Original post →

More from Infra

Infra channel →