llama.cpp benchmarks show ROCm and Vulkan trading wins on AMD Radeon AI PRO R9700

Gesha24 · reddit · 2026-07-28

A llama.cpp user benchmarked ROCm 7.14 against Vulkan on an AMD Radeon AI PRO R9700 and found backend-dependent results that are not as obvious as they first look.

The post includes full hardware, build, and model details, then compares short-context and long-context runs across E4B-Q4, Gemma 31B Q4, and Qwen 27B Q6. In short-context tests, ROCm is faster on some prompt-processing runs, while Vulkan wins on others and sometimes edges out ROCm on token generation. In the longer-context runs, the gap shifts again, with both backends trading wins depending on model and metric.

The takeaway is that on RDNA4 hardware, backend choice in llama.cpp can materially change throughput, but the direction of the advantage depends on the model size and context pattern rather than being universally better for one stack.

Original post →

More from Infra

Infra channel →