85.6 tok/s on a single RTX 5090: local Qwen3 27B with MTP speedup

comperr · reddit · 2026-09-26

A Reddit user reports running Qwen3.8:27B locally on a single RTX 5090, achieving 85.6 tok/s using MTP (multi-token prediction). A useful data point for anyone considering running 27B-class models on consumer hardware.

Original post →

More from Infra

Infra channel →