Running Qwen 27B on a Single 5090: 200+ t/s Inference Benchmarks
Maleficent-Ad5999 · reddit · 2026-08-31
A user successfully ran the Qwen3.8-27B model (NVFP4 quant) on a single RTX 5090 (32GB), achieving 200+ t/s decode speeds using the ninfer engine. Benchmarks show that with 180K context and MTP (5 draft tokens) enabled, decode speed reached 200.5 tok/s. The author provides specific launch commands and benchmark results, noting that while synthetic text is easy to predict (94-100% accept), real agentic coding scenarios see 51% accept rates, dropping speed to 154 tok/s.
More from Infra
- SK Hynix breaks ground on Indiana HBM plant, targeting HBM4e mass production in 2029 — Beth_Kindig · 2026-08-31
- Qwen 3.8 Flash Next runs at 3.5 tok/s on mid-range Android phone — dai_app · 2026-08-31
- The Boring Company's sales dilemma: 20x cost advantage but no sales team — PTrubey · 2026-08-31
- Running 182B Qwen on 4080: 8 tok/s via SSD offloading — Desperate-Data-3747 · 2026-08-31
- AI is transforming how energy systems are monitored, maintained and operated — ingliguori · 2026-08-31
- Uber Cuts AI Costs 52% While 10xing Usage via Agentic Workflow — alvelda · 2026-08-31