$140 Radeon MI50 paired with GTX-1080Ti boosts local 27B-35B LLM speeds up to 9x
tabletuser_blogspot · reddit · 2026-09-20
A Reddit user added a used AMD Radeon MI50 (16GB, $140, firmware-flashed to Radeon VII) to a GTX-1080Ti rig for 27GB total VRAM and benchmarked 30B-class local models with llama.cpp's Vulkan build on Kubuntu.
- Tested Qwen3.5-35B MoE (MXFP4), Nemotron-31B MoE, Qwen3.5-27B, Gemma3-27B and Gemma4-26B MoE GGUF quants, comparing single vs dual GPU with the GGMLVKVISIBLEDEVICES flag and full llama-bench commands included
- Dense models gained the most: Gemma3 27B went from 1.5 t/s to 14.96 t/s (896%), Qwen3.5 27B from 2.1 to 15.6 t/s (+642%), Nemotron 31B +504%
- No driver conflicts thanks to Vulkan; MoE models were already fast on one card
- Takeaway: for under $140, running 30B dense models locally becomes a practical option
More from Infra
- Qwen3.8 Flash Next on one RTX 5090 hits 50 t/s decode via FreeToken expert caching — dir3ctly · 2026-09-20
- Are data centers dodging taxes? Tax Foundation data on $1B facilities says no — AndyMasley · 2026-09-20
- AMD carries out serious software optimizations for Kimi-K3, analyst says — AccBalanced · 2026-09-20
- RAM price jumps from $350 to $640 in months, pricing out new PCs — chrisalbon · 2026-09-20
- NVIDIA engineer breaks down why DeepSeek re-engineered V4.1 Flash for speed — thursdai_pod · 2026-09-20
- halogen 0.12.0 hits 38 tok/s decode at 1M context for Qwen on Strix Halo — peonist-ai · 2026-09-20