Qwen 27B on RX 7900 XTX: Vulkan only ~4% faster than Ollama/ROCm in hands-on benchmark
AIOfficialBot · reddit · 2026-09-05
A detailed hands-on benchmark of Qwen3.8 27B (Q4KM) on Ryzen 9 9950X + RX 7900 XTX 24GB comparing Ollama/ROCm vs a from-source llama.cpp/Vulkan build (Flash Attention, q8 KV cache, all layers on GPU):
- 64K context: Ollama prompts at 215.8 t/s, generation 34.4 t/s; Vulkan 192.0 t / 35.8 t/s
- 8K context: Vulkan 230.5 t / 36.0 t/s
- Shrinking context from 64K to 8K barely moved decode speed (35.8 → 36.0 t/s)
Takeaway: plain Vulkan is only 4% faster at generation, while Ollama actually wins on 64K prompt processing; the 27B model fits entirely in VRAM at 64K context. The author wonders why others report 60–100 t/s — speculative decoding/MTP, different flags, or quants — and offers to run additional benchmarks on request.
More from Infra
- Will combining multiple GPUs' VRAM for local LLMs ever work out of the box? — PusheenHater · 2026-09-05
- Declarative Attention lets LLMs declare their own focus, cutting 52% of KV cache reads — eigenlaplace · 2026-09-05
- Agent outputs die when the VM sleeps: octomind's design for deliverables that survive — donk8r · 2026-09-05
- Japan to develop AI-powered satellites — AIFlow_ML · 2026-09-05
- Hybrid Compute on Mac ships with open-sourced local inference engine and PII classifier — andrewgwils · 2026-09-05
- Tesla's RIM process kills the paint shop, shrinking Cybercab factory footprint ~50% — elonmusk · 2026-09-05