Mixed old GPUs + 32GB RAM runs Qwen3 27B at 20t/s for local coding
pepijndevos · reddit · 2026-09-01
The author wedged an AMD RX 7900 and an Nvidia Quadro 5000 into a 32GB RAM workstation using a riser and Lego, running Qwen3 27B q4km with 128k context fully GPU-resident via llama.cpp at 20 tokens/s. He finds it totally viable as a local coding model, just slow — and asks the community for a $1k upgrade (system RAM for MoE models, or a less mismatched GPU setup) instead of a $10k machine.
More from coding & agent
- SparkLLM releases on-device models with 1M context support — MaziyarPanahi · 2026-09-01
- Scheduling swarm agents for poetry/art competitions during morning routine — cephaloform · 2026-09-01
- Building personal agents: insanely cheap, always-on, and personality-adjusting — bindureddy · 2026-09-01
- Introducing Loupe: An open-source PR review agent — andersonbcdefg · 2026-09-01
- Why "it feels better" isn't good enough for production LLM decisions — camerongreen95 · 2026-09-01
- Hermes Agent monitors chats and emails to auto-create tasks and draft work — intellectronica · 2026-09-01