AMD V620 benchmarks: Qwen 27B vs Gemma A4B on llama.cpp

Brave_Load7620 · reddit · 2026-08-22

A developer shared benchmarks running llama.cpp on an AMD V620 GPU via Vulkan and ROCm on Windows 11. The tests compared Qwen 27B and Gemma A4B across different context depths (3.4k to 26.6k tokens), measuring prefill (PP) and generation speeds. The results show that the Gemma model performs significantly better on the ROCm backend, with generation speeds far surpassing Qwen, especially in long-context scenarios. The author is still optimizing flags and settings.

Original post →

More from Infra

Infra channel →