Qwen3.8-27B hits 223 tok/s on RTX 6000 Pro with NVFP4
BanghuaZ · x · 2026-08-17
This repository provides scripts to run Qwen3.8-27B via SGLang on NVIDIA RTX 6000 Pro (96GB). Utilizing NVFP4 weights and DSpark speculative decoding, it achieves 200-223 tok/s single-stream throughput. The setup supports the native 256k context window and 8 concurrent requests by default, using FP8 KV cache for optimal memory usage.
More from Infra
- AI’s next bottleneck is power: The race for energy — ingliguori · 2026-08-17
- System design in 2026: Master these 10 core concepts — blaizedsouza · 2026-08-17
- AI Economics: Why compute and power are the new bottlenecks — demian_ai · 2026-08-17
- Building production-grade service layers for agents — blaizedsouza · 2026-08-17
- Expert Concerns Embedded Devices as a Weak Link in AI Security — johnowhitaker · 2026-08-17
- Help: Converting Qwen3.8 27B to ONNX for NPU Usage — xXDennisXx3000 · 2026-08-17