Pairing a decade-old RX 480 with RX 7900 GRE boosts llama.cpp MoE inference 36% over single GPU
tabletuser_blogspot · reddit · 2026-09-08
A Reddit user benchmarked dual-GPU llama.cpp inference (Vulkan build 10453) with an RX 7900 GRE 16GB plus an old RX 480 8GB.
- Gemma4 26B-A4B Q4KM: 52.03 t/s generation on two GPUs vs 37.93 t/s on the 7900 GRE alone (36% faster); prefill rose from 238 to 321.6 t/s.
- Dense 27B models crawl on this setup: Qwen35 27B Q6K at 5.96 t/s, Q5KM at 11.73 t/s.
- Flash Attention made no measurable difference for the MoE model (52.03 vs 51.94 t/s).
- Even the low-bandwidth RX 480 (256 GB/s, 1.8 t/s alone) meaningfully helps MoE inference when pooled.
More from Infra
- Jensen Huang confirms GPT-6 Astra trained on 100K+ Grace Blackwell NVL72 systems — rohanpaul_ai · 2026-09-08
- South Korea to give everyone free generative AI, backed by up to 512 B200 GPUs — IgorCarron · 2026-09-08
- Dev runs SDXL fine-tune fully on iPhone Neural Engine: 6-bit, 8 steps, offline — NovaDevCodeStudio · 2026-09-08
- Qwen 27B q8 vs bf16 on a DGX Spark: is the 1% token difference worth the memory? — superSmitty9999 · 2026-09-08
- Running Qwen3.8-27B for coding on 32GB VRAM — what local LLMs do you use and why? — theexile1337 · 2026-09-08
- A simple Windows tray app for monitoring NVIDIA GPU VRAM — Due-Committee-9591 · 2026-09-08