Qwen3.8-27B Benchmarks: 672 tps on RTX 3090

Local benchmarks of Qwen3.8-27B have poured in since its release: SGLang offers Day-0 support, and with NVFP4 quantization plus DSpark speculative decoding, a single RTX 5090 decodes at 206.1 tok/s while the RTX 6000 Pro (96GB) hits 200-223 tok/s in single-stream throughput; the RTX 4090 delivers 73 tok/s for coding with a 128k context, which testers deem fully usable. A 27B-class model now runs smoothly on a single consumer GPU.

Confirmed

Unconfirmed

Why it matters

2026-08-16 ~ 2026-08-17 · 6 related posts

Primary sources