6x BC-250 mining board cluster runs local LLMs at 100k context, 28 tok/s
Ok-Breadfruit-3523 · reddit · 2026-10-02
A Redditor built a local LLM cluster from 6 BC-250 ex-mining boards in an ASRock 4U12G case: 4 boards run Qwen Next Flash IQ2XS at 100k context (28 tok/s, 115 ppt at 50k), while the other two run Qwen 3.6 35B Q4 at 60 tok/s with 100k context and 450 ppt. Everything runs llama.cpp with Vulkan over 1Gb Ethernet RPC — with a cardboard box topped by 3 fans as intake.
More from Infra
- awesome-jev indexes 700 production tools around TypeSafe AI's decision model Jev — Remarkable-Gur719 · 2026-10-02
- smolvm v1.22 ships near-instant VM resume for undoing agent actions, 6.5k stars — LoganGrasby · 2026-10-02
- Broadcom to lend Anthropic up to $42B for AI chips, eyeing top customer slot by 2027 — rohanpaul_ai · 2026-10-02
- Meta paper: only 50-60% of recommendation training time actually trained before optimizations — _reachsumit · 2026-10-02
- Dev open-sources GPT-2-tools to run original 1.5B GPT-2 XL locally on CPU — MikePFrank · 2026-10-02
- CoreWeave launches serverless GPUs: hourly-billed, no contract, private preview — altryne · 2026-10-02