RTX 5000 Pro Runs Qwen3.8 at Unbeatable Value
Valuable-Run2129 · reddit · 2026-08-24
A user shared impressive benchmarks running Qwen3.8-27B on an RTX 5000 Pro (48GB VRAM): 130 t/s decode, 4000 t/s prefill, and 5 concurrent requests. Using SGLang, prefixes are offloaded to storage to prevent reprocessing in sub-agentic workflows. The author argues this model drastically increases local compute value, rivaling a Claude subscription and driving up the card's resale price.
More from Infra
- Context Layer: A User-Controlled Contract for Cross-Model/Agent Context Transfer — sierracatalina · 2026-08-24
- OpenRouter: token demand will far outpace compute; heterogeneous hardware including retail Macs is key — gajesh · 2026-08-24
- Decentralized Inference on MacBooks: Earn $200-$500/Month via darkbloom — siddharthnibjiya · 2026-08-24
- Google Colab adds SSH and CLI support for seamless cloud workflows — HankYeomans · 2026-08-24
- Speculative Decoding: Obsolete by the Time You Implement It — dejavucoder · 2026-08-24
- EigenLabs Serves 5.23B Tokens in a Day, Hits $134K ARR — AccBalanced · 2026-08-24