DIY CUDA box with unlocked CMP170HX mining cards hits 4000 tps prompt processing for Qwen Flash Next
Miserable-Dare5090 · reddit · 2026-09-02
A Reddit user shared a full build log for a DIY CUDA server running LLMs on two unlocked CMP170HX mining cards.
- Expanded PCIe lanes via cheap Newegg CPU/MB combos; reused 64GB DDR5 and SSDs from an old prebuilt
- Budget case + Noctua fans; both CMP170s in a GPU switch on the main x16 slot, other cards on oculink
- Qwen Flash Next fits nicely: 4000 tps prompt processing, 80+ tps single-stream decode without MTP
- DeepSeek repo would need at least 3x 64GB cards — risky as mining card prices rise
- Now runs three models simultaneously: Flash Next, Qwen 27B, Gemma 26B
More from Infra
- Over 70% of Voters Dislike Data Centers as Opposition Swells — WillRinehart · 2026-09-02
- AMD MI355X sees major performance gains on agentic workloads — AccBalanced · 2026-09-02
- Dell posts record $47B quarter as AI server revenue doubles to $16.4B — Beth_Kindig · 2026-09-02
- Alchemy PR adds native Cloudflare tracing via a single Telemetry layer — samgoodwin89 · 2026-09-02
- Wall Street's Dell estimates were off by a light year amid Nvidia AI server supercycle — firstadopter · 2026-09-02
- Distributed Dev Setup: Seamless Switching Without Restarts — doodlestein · 2026-09-02