$250 of modded mining cards, 30GB VRAM: old i7 PC runs Qwen at 30 tok/s with patched drivers
HFq_Dev · reddit · 2026-09-19
A Reddit user turned an old i7-4790K/Z97 PC into a local inference server after his GTX 1070 died during maintenance. Instead of buying mainstream GPUs, he grabbed two modded CMP50HX mining cards (one 20GB with PCIe x16 mod, one stock 10GB) for $250 total — prices have since doubled.
The magic is community driver patches that nearly remove the cards' compute limits and add PCIe 2.0 support (natively PCIe 1.1), roughly doubling performance: Qwen3.8 27B with MTP runs at 30-35 tok/s with 300-400 tok/s prompt processing, and MoE model Ornith 1.5 35BA3B hits 80-100 tok/s. He caps power at 180W due to his PSU (225W would be faster).
Gotchas documented: early patches didn't support the 20GB variant, and sleep/wake crashes the driver (fixed by disabling sleep; traced to the PCIe patch). He maintains multiple presets for different quantizations and KV cache configs — a full playbook for budget local deployment.
More from Infra
- First Cafe Compute Nairobi meetup demos Cerebras API, calls for an /explain interpretability endpoint — paw_lean · 2026-09-19
- Cerebras launches Money Agent, a Qwen 3 27B-powered personal finance assistant — Alibaba_Qwen · 2026-09-19
- GPU host warns: renter exploited his rig for attacks, Clore.AI blocked him for reporting it — anomaly256 · 2026-09-19
- 3M paid $10.3B to quit PFAS — AI data centers just made it a growth market again — aakashgupta · 2026-09-19
- Random KV Cache eviction rivals top baselines and boosts vLLM throughput 32-43% — 机器之心 · 2026-09-19
- Reading a Pretraining Run: A P0/P1/P2 Metric System for Monitoring LLM Pretraining — SonglinYang4 · 2026-09-19