llama.cpp on 2x RTX 2060 hits 45 tok/s on Qwen3.8 27B — is dual RX 6800 worth it?

BarberIcy366 · reddit · 2026-09-09

A Reddit user shares real llama.cpp numbers: two RTX 2060 12GB cards (24GB total) running Qwen3.8 27B IQ4XS at 131k context get 45 tok/s. Eyeing a deal on RX 6800 + 6800 XT (32GB), they ask whether higher bandwidth and VRAM actually translate to gains in mixed AMD multi-GPU llama.cpp setups, and how ROCm/Vulkan support on the 6800 series holds up — useful reference data for budget local inference.

Original post →

More from Infra

Infra channel →