Running Qwen 27B on a Single 5090: 200+ t/s Inference Benchmarks

Maleficent-Ad5999 · reddit · 2026-08-31

A user successfully ran the Qwen3.8-27B model (NVFP4 quant) on a single RTX 5090 (32GB), achieving 200+ t/s decode speeds using the ninfer engine. Benchmarks show that with 180K context and MTP (5 draft tokens) enabled, decode speed reached 200.5 tok/s. The author provides specific launch commands and benchmark results, noting that while synthetic text is easy to predict (94-100% accept), real agentic coding scenarios see 51% accept rates, dropping speed to 154 tok/s.

Original post →

More from Infra

Infra channel →