Mixed old GPUs + 32GB RAM runs Qwen3 27B at 20t/s for local coding

pepijndevos · reddit · 2026-09-01

The author wedged an AMD RX 7900 and an Nvidia Quadro 5000 into a 32GB RAM workstation using a riser and Lego, running Qwen3 27B q4km with 128k context fully GPU-resident via llama.cpp at 20 tokens/s. He finds it totally viable as a local coding model, just slow — and asks the community for a $1k upgrade (system RAM for MoE models, or a less mismatched GPU setup) instead of a $10k machine.

Original post →

More from coding & agent

coding & agent channel →