Qwen3.8-27B hits 223 tok/s on RTX 6000 Pro with NVFP4

BanghuaZ · x · 2026-08-17

This repository provides scripts to run Qwen3.8-27B via SGLang on NVIDIA RTX 6000 Pro (96GB). Utilizing NVFP4 weights and DSpark speculative decoding, it achieves 200-223 tok/s single-stream throughput. The setup supports the native 256k context window and 8 concurrent requests by default, using FP8 KV cache for optimal memory usage.

Original post →

More from Infra

Infra channel →