$135 MI50 + RX 7900 GRE dual-GPU Vulkan benchmarks across 19 GGUF models
tabletuser_blogspot · reddit · 2026-09-19
The author bought a China-market MI50 (Radeon VII, 16GB) for about $135, paired it with an RX 7900 GRE for 32GB total VRAM, and ran llama.cpp's Ubuntu Vulkan prebuilt binary (power-limited to 220/190W) across 19 GGUF models from 20B to 35B.
Highlights:
- Fastest overall: gpt-oss 20B Q6K (pp512 660 t/s, tg128 84 t/s)
- MoE models shine: qwen3moe 30B Q6K at 67 t/s tg128, cohere2moe 30B MXFP4 at 504 t/s pp512
- Dense 27–31B models are much slower (gemma4 31B Q6K only 12.6 t/s)
- Full tables sorted both by model and by parameter count — a useful reference for budget dual-GPU inference
More from Infra
- VanEck: NVDA's bigger risk is customers can't get power; powered-land base case implies ~83% upside — menhguin · 2026-09-19
- Reverse-engineering Claude's subscription limits from unrounded floats: Max 5× is the real sweet spot — RexDouglass · 2026-09-19
- Hyperscaler ROIIC Peaked Near 40% vs 8% Cost of Capital, AI Capex Math Shows — menhguin · 2026-09-19
- ByteDance said to dominate Malaysia datacenter capacity as China buildout sparks debate — jwt0625 · 2026-09-19
- Databricks brings Unity Gateway to Neon, billed as fastest AI gateway for Kimi K3 — Yuchenj_UW · 2026-09-19
- Local Qwen matches Jev at 96.53% accuracy, 239 ms vs 368 ms median latency — Ok-Development6070 · 2026-09-19