$100 of used RX 580s runs Qwen 27B at 7.39 t/s on DDR3 platform
Whole_Alternative_18 · reddit · 2026-08-17
A Reddit user built a $100 rig from two used RX 580 8GB cards (16GB VRAM total) on an old DDR3 workstation board and runs Qwen 27B at Q3KM quantization, getting 7.4 t/s generation (14 t/s processing).
They admit it's far from comfortable: slow speed, painful long-context input, and unsuitable for MTP-style models. But it's extremely cheap—even cheaper than running on system RAM—making it a viable last-resort option for tight budgets.
More from Infra
- Qwen 3.8 27B runs at only 20t/s on AMD 7900XTX GPU — soyalemujica · 2026-08-17
- omlx: LLM Inference Server with SSD Caching for Apple Silicon — jundot · 2026-08-17
- Dev open-sources orchestration system that lets local 27B models pick stacks and ship production code — Single_Land8080 · 2026-08-17
- Ajinomoto cuts supply 30%, threatening China's AI supply chain — pstAsiatech · 2026-08-17
- Cerebras accelerates RL inference: 10-hour tasks finish in 1 hour — dejavucoder · 2026-08-17
- llmfit scans your hardware to find the optimal LLMs you can actually run — mhdfaran · 2026-08-17