$100 used Tesla P100 beats gaming GPU nearly 2x for local LLM inference
Mrinohk · reddit · 2026-09-28
A local inference enthusiast shared a hands-on test that contradicted Claude's advice against used Tesla P100s (no tensor cores, weak quantization, power/cooling hassle). After buying one for under $100 and applying community llama.cpp patches with --cpu-moe, results far exceeded expectations.
Key numbers: on an RX6600 XT, Qwen3.6 35B A3B (Q4 + MTP) ran at 30-35 tok/s generation (45-50 for code); the P100 with shinbunbun's dedicated patches hit 54-60 tok/s (66-72 for code) at 32k context. Cooling uses a 3D-printed housing with a 94mm Noctua fan; a new PSU cost about $100.
Takeaway: don't fully trust frontier models on budget hardware. The P100 is a seriously underrated entry card for local inference—nothing else comes close at that price. A second card is on the way for a dedicated dual-GPU inference box.
More from Infra
- Debate: An AI That Only Speaks PTX — Kernel Optimization Without Natural Language — teortaxesTex · 2026-09-28
- TrendForce sees 268GW data center power gap by 2030, US alone over 170GW — Beth_Kindig · 2026-09-28
- Dell up 338% YTD as AI compute stocks deliver huge gains: MU, INTC, AMD — firstadopter · 2026-09-28
- Nvidia targets 500,000 RTX Pro 5500 chips per quarter for China; ByteDance order alone takes two quarters — pstAsiatech · 2026-09-28
- Nvidia says China's chip makers saw 'unprecedented growth' since 2022 amid outdated US export controls — pstAsiatech · 2026-09-28
- From 16GB to 128GB: A Local AI Hardware Journey — TheOyinbooke · 2026-09-28