Qwen 3.8 27B runs at only 20t/s on AMD 7900XTX GPU
soyalemujica · reddit · 2026-08-17
A user reported extremely slow inference speeds (20-35 t/s) when running Qwen 3.8 27B on an AMD 7900XTX, even with Speculative MTP enabled at long contexts (80k+). Detailed launch arguments are provided, suspecting architecture or quantization issues.
More from Infra
- Local AI Primer: Benchmarking 2B to 27B Models on Consumer Hardware — draginol · 2026-08-17
- LocalAI adds 17 C++ inference backends in six months — rms80 · 2026-08-17
- Nvidia: Land, Power, and Shell Are Critical Resources for AI Factories — firstadopter · 2026-08-17
- Combining 3060 and 4070 Super for local AI: feasibility and setup — HsSekhon · 2026-08-17
- NVIDIA Invests $1.5B to Secure 8GW Capacity for OpenAI in Ohio — nvidia · 2026-08-17
- NVIDIA to Provide Exclusive AI Infrastructure at Ohio Campus for OpenAI — nvidia · 2026-08-17