llama.cpp on 2x RTX 2060 hits 45 tok/s on Qwen3.8 27B — is dual RX 6800 worth it?
BarberIcy366 · reddit · 2026-09-09
A Reddit user shares real llama.cpp numbers: two RTX 2060 12GB cards (24GB total) running Qwen3.8 27B IQ4XS at 131k context get 45 tok/s. Eyeing a deal on RX 6800 + 6800 XT (32GB), they ask whether higher bandwidth and VRAM actually translate to gains in mixed AMD multi-GPU llama.cpp setups, and how ROCm/Vulkan support on the 6800 series holds up — useful reference data for budget local inference.
More from Infra
- Proteus Generates Custom GPU Kernels for Qwen3 122B, Up to 5.2x Faster Than vLLM — matei_zaharia · 2026-09-09
- Google Cloud August AI Infra Roundup: gVisor Sandboxes on Ray, Filestore on Colossus — dl_weekly · 2026-09-09
- Provably private inference services promise prompts unreadable to providers — corbtt · 2026-09-09
- 300B output tokens, $20-30M in compute: Ethan Mollick says AI science will need far more compute — eldonredwards · 2026-09-09
- Rumor resurfaces: Google may replace Nvidia as TSMC's biggest customer — zephyr_z9 · 2026-09-09
- Tahuna open-sources ephemeral GPU orchestration for ML workloads — Monaim101 · 2026-09-09