6x BC-250 mining board cluster runs local LLMs at 100k context, 28 tok/s

Ok-Breadfruit-3523 · reddit · 2026-10-02

A Redditor built a local LLM cluster from 6 BC-250 ex-mining boards in an ASRock 4U12G case: 4 boards run Qwen Next Flash IQ2XS at 100k context (28 tok/s, 115 ppt at 50k), while the other two run Qwen 3.6 35B Q4 at 60 tok/s with 100k context and 450 ppt. Everything runs llama.cpp with Vulkan over 1Gb Ethernet RPC — with a cardboard box topped by 3 fans as intake.

Original post →

More from Infra

Infra channel →