$500 ex-mining BC-250 cluster runs Qwen 35B at 145 tok/s with 256k context
Ok-Breadfruit-3523 · reddit · 2026-10-11
A local AI hobbyist's third update on a DIY cluster: 4 used BC-250 mining boards (<$100 each, now $150) linked via 2.5GbE, running llama with Vulkan + RPC, achieving 60-70 tok/s for Gemini Next Flash IQ3XXS at 100k context, or two Qwen 3.6 35B A3B IQ4 instances at 145 tok/s each with usable context extended to 258k.
Iterating with Claude Code, key optimizations include: chained speculative decoding with the model's own MTP draft head (Q20 from 61 to 68 tok/s), recording and replaying Vulkan command graphs (saving 3.7-4.5ms per graph), custom kernels for the architecture (+7.3% at 62k context), copying context checkpoints on-board instead of via the coordinator (665ms→18ms, prompt speed +60%), and a per-token attention mask limit (saving 768MiB).
Power draw is 600-900W under load. Next steps: AIO coolers per board to fix thermal throttling and a 3D-printed cooling lid replacing the cardboard plenum. The bottleneck is latency, not bandwidth, so faster networking wouldn't help.
More from coding & agent
- Zed CEO's dashboard shows 295M tokens burned by 2 devs in 30 days, mostly cache reads — zeeg · 2026-10-11
- Claude Code's TAB+ENTER flow: accepting recommended next steps for semi-autonomous coding — tristanbob · 2026-10-11
- banteg echoes DHH talk: knowing less helps, since overengineered prompts hold you back — banteg · 2026-10-11
- LLVM Internals Series: Writing a Simple LLVM Pass From Scratch — sh4dy_0011 · 2026-10-11
- Beff Jezos: The opportunity cost of not vibe coding on weekend evenings is too high — beffjezos · 2026-10-11
- Dev ditches Codex for game UI work after finding Claude Code output looks far better — tristanbob · 2026-10-11