Qwen 2.5 27B Runs at 100+ tok/s on RTX 5090
Hesamation · x · 2026-08-18
A demonstration shows Qwen 2.5 (mistakenly labeled Qwen 3.8 in the tweet) 27B model running locally on an NVIDIA RTX 5090 gaming GPU, achieving over 100 tokens per second. This highlights how frontier AI performance from six months ago is now accessible on consumer hardware.
More from Infra
- Is FP8 Quant a Bad Idea for Qwen 3.8 27B? Quality Concerns Raised — lblblllb · 2026-08-18
- Unsloth's Qwen 3.8 Q4KXL Fails in Long Context, Bartowski Holds Up — Healthy-Nebula-3603 · 2026-08-18
- Infrastructure Unlocking Deep Network Power: Speed of Build and Test is a Massive Differentiator — chris_j_paxton · 2026-08-18
- AI power thesis turns real: utilities CEG, NEE, Vistra monetize data center demand — Efficient_Ad5893 · 2026-08-18
- StringZillas CUDA achieves 32 GB/s multi-pattern matching on single GPU — srchvrs · 2026-08-18
- Asking for help optimizing AMD 7800 XT for Krea 2 generation — TheHumanCrab · 2026-08-18