AMD V620 benchmarks: Qwen 27B vs Gemma A4B on llama.cpp
Brave_Load7620 · reddit · 2026-08-22
A developer shared benchmarks running llama.cpp on an AMD V620 GPU via Vulkan and ROCm on Windows 11. The tests compared Qwen 27B and Gemma A4B across different context depths (3.4k to 26.6k tokens), measuring prefill (PP) and generation speeds. The results show that the Gemma model performs significantly better on the ROCm backend, with generation speeds far surpassing Qwen, especially in long-context scenarios. The author is still optimizing flags and settings.
More from Infra
- Seeking benchmarks: 4x DGX Spark cluster vs 2x cards — Gobra_Slo · 2026-08-23
- Dual RTX 3060 12GB performance for Qwen3.8-27B — Mean-Ad1493 · 2026-08-23
- Local GPU-powered home industrial machines: the overlooked AI hardware frontier — curious_vii · 2026-08-23
- Tests show Quantization has minimal impact until below Q4 for local LLMs — KitchenAmoeba4438 · 2026-08-23
- Open Source AI Token Share on Vercel Surges to 62% — GavinSBaker · 2026-08-23
- Engineers Debate PIC Design: Foundry Incompetence and Lagging COUPE — jwt0625 · 2026-08-23