$500 ex-mining BC-250 cluster runs Qwen 35B at 145 tok/s with 256k context

Ok-Breadfruit-3523 · reddit · 2026-10-11

A local AI hobbyist's third update on a DIY cluster: 4 used BC-250 mining boards (<$100 each, now $150) linked via 2.5GbE, running llama with Vulkan + RPC, achieving 60-70 tok/s for Gemini Next Flash IQ3XXS at 100k context, or two Qwen 3.6 35B A3B IQ4 instances at 145 tok/s each with usable context extended to 258k.

Iterating with Claude Code, key optimizations include: chained speculative decoding with the model's own MTP draft head (Q20 from 61 to 68 tok/s), recording and replaying Vulkan command graphs (saving 3.7-4.5ms per graph), custom kernels for the architecture (+7.3% at 62k context), copying context checkpoints on-board instead of via the coordinator (665ms→18ms, prompt speed +60%), and a per-token attention mask limit (saving 768MiB).

Power draw is 600-900W under load. Next steps: AIO coolers per board to fix thermal throttling and a 3D-printed cooling lid replacing the cardboard plenum. The bottleneck is latency, not bandwidth, so faster networking wouldn't help.

Original post →

More from coding & agent

coding & agent channel →