Two $115 ex-mining APUs run Qwen 35B at 60 tok/s with 64K context for ~$300
Ok-Breadfruit-3523 · reddit · 2026-09-24
A Reddit user built a budget local LLM rig from two ex-mining BC-250 APUs ($115 each, 27GB combined GPU memory), linked via llama.cpp with Vulkan and RPC over 1Gb Ethernet on Bazzite. The setup runs Qwen3.6-35B-A3B at Q4KM, reaching 60 tok/s with 64K context — roughly $300 total including PSU. The author plans to expand to six boards to try running Qwen 3.8 Flash. A replicable reference for budget local deployment.
More from Infra
- Alibaba's Banma ships AutoOmni 2.0: 3B-active edge model nears 10x-larger cloud models — 机器之心 · 2026-09-24
- Huang and Musk: China reaches advanced lithography in 2-3 years; chip bans only buy time — beffjezos · 2026-09-24
- MacBook Pro M5 Max runs AMD Radeon AI PRO R9700 over Thunderbolt 5 for local LLMs — TheOriginalG2 · 2026-09-24
- A real-time voice agent with zero US servers: Gladia STT, Gemma 4 on Scaleway, KugelAudio TTS — tobowers · 2026-09-24
- Running MiniCPM5-2B as a local agent on M4 16GB: trimming, tool-call adapter and benchmarks — 面壁智能 · 2026-09-24
- WSJ: The AI build-out is becoming the biggest economic bet in U.S. history — GeneReddit123 · 2026-09-24