$800 rig of 5 ex-mining BC-250 boards runs Qwen3-Coder-Next at 40 tok/s
Ok-Breadfruit-3523 · reddit · 2026-09-28
A Reddit user built a local inference rig from five retired BC-250 mining boards for under $800, exposing 71GB of VRAM across an Asrock 12-unit case with one head unit and the rest headless over 1Gb Ethernet. Qwen3-Coder-Next Q4 hits 40 tok/s at 30k context, dipping to 30 tok/s at 100k. Power efficiency is terrible, but the cost-to-performance is remarkable; two more boards are queued for testing 3.8 flash next.
More from Infra
- Chamath breaks down AI compute: prefill is compute-bound, decode is bandwidth-bound — rohanpaul_ai · 2026-09-28
- Taiwan companies scramble for advanced packaging talent, says analyst — LIWEI_TWCapital · 2026-09-28
- Developer says local AI is shifting from nice-to-have to infrastructure: control beats privacy — Aiden_Tech_Ai · 2026-09-28
- Meta open-sources Component Benchmark, a hierarchical profiler for TB-scale recommender models — _reachsumit · 2026-09-28
- apple-llm: Node/Python wrapper for the free local LLM built into Apple Silicon Macs — light_2earth · 2026-09-28
- Running 8 watercooled GPUs for local AI: one user's case for watercooling over air cooling — HanchungLee · 2026-09-28