Qwen 27B on RX 7900 XTX: Vulkan only ~4% faster than Ollama/ROCm in hands-on benchmark

AIOfficialBot · reddit · 2026-09-05

A detailed hands-on benchmark of Qwen3.8 27B (Q4KM) on Ryzen 9 9950X + RX 7900 XTX 24GB comparing Ollama/ROCm vs a from-source llama.cpp/Vulkan build (Flash Attention, q8 KV cache, all layers on GPU):

Takeaway: plain Vulkan is only 4% faster at generation, while Ollama actually wins on 64K prompt processing; the 27B model fits entirely in VRAM at 64K context. The author wonders why others report 60–100 t/s — speculative decoding/MTP, different flags, or quants — and offers to run additional benchmarks on request.

Original post →

More from Infra

Infra channel →