llama.cpp benchmarks show ROCm and Vulkan trading wins on AMD Radeon AI PRO R9700
Gesha24 · reddit · 2026-07-28
A llama.cpp user benchmarked ROCm 7.14 against Vulkan on an AMD Radeon AI PRO R9700 and found backend-dependent results that are not as obvious as they first look.
The post includes full hardware, build, and model details, then compares short-context and long-context runs across E4B-Q4, Gemma 31B Q4, and Qwen 27B Q6. In short-context tests, ROCm is faster on some prompt-processing runs, while Vulkan wins on others and sometimes edges out ROCm on token generation. In the longer-context runs, the gap shifts again, with both backends trading wins depending on model and metric.
The takeaway is that on RDNA4 hardware, backend choice in llama.cpp can materially change throughput, but the direction of the advantage depends on the model size and context pattern rather than being universally better for one stack.
More from Infra
- Kimi K3 runs on AMD MI350X with SGLang and hits 327 tok/s across four requests — burny_tech · 2026-07-28
- Kimi K3 reportedly keeps block attention residuals in a 93-layer design — burny_tech · 2026-07-28
- NASA’s new administrator says SpaceX-backed orbital data centers will happen — elonmusk · 2026-07-28
- Reddit users test whether unlocked CMP 170HX cards can power local AI rigs — Tritheone69 · 2026-07-28
- Open-weight models are pitched as a security win for large companies — jessi_cata · 2026-07-28
- Dolphin is being trained on Trinity Large Thinking 398B across 72 RTX 4090s — QuixiAI · 2026-07-28