Qwen 2.5 27B Runs at 100+ tok/s on RTX 5090

Hesamation · x · 2026-08-18

A demonstration shows Qwen 2.5 (mistakenly labeled Qwen 3.8 in the tweet) 27B model running locally on an NVIDIA RTX 5090 gaming GPU, achieving over 100 tokens per second. This highlights how frontier AI performance from six months ago is now accessible on consumer hardware.

Original post →

More from Infra

Infra channel →