Open-source QuantProbe predicts local-LLM speed on a 2016 PC and a GTX 1060
Ok_Brush_3449 · reddit · 2026-07-23
QuantProbe is an open-source project for choosing quantization and memory-allocation recipes to run local LLMs on your own machine.
The author reports that a 110B model on a 2016 PC with 16GB RAM and SATA storage was predicted to run at 0.2–0.3 tok/s and measured at 0.19 tok/s, while the same “law” yielded 19.3 tok/s for a 30B model on a GTX 1060 6GB. The goal is to make local-model performance more predictable and accessible.
The project is looking for testers and contributors, and the author says the work could become a starting point for further research on token economy and local inference tradeoffs.
More from Infra
- Science paper turns memristor drift into a 2.12 ms neural dynamical system — jiqizhixin · 2026-07-23
- Proprio Robotics pitches robots for autonomous data centers as AI strains the grid — ycombinator · 2026-07-23
- AMD Showcases Large-Scale MoE Serving on Instinct MI355X at Advancing AI — ying11231 · 2026-07-23
- Korea's NAND Flash Exports Surge 300% YoY in June Driven by AI Inference Demand — tengyanAI · 2026-07-23
- Open-Sourced DSpark Speculator: Boosts Inkling Decode Throughput by 1.89X — BanghuaZ · 2026-07-23
- Nvidia-backed Fireworks AI raises $1.5B at $17.5B valuation, ARR tops $1B, daily tokens hit 40 trillion — Beth_Kindig · 2026-07-23