Open-source QuantProbe predicts local-LLM speed on a 2016 PC and a GTX 1060

Ok_Brush_3449 · reddit · 2026-07-23

QuantProbe is an open-source project for choosing quantization and memory-allocation recipes to run local LLMs on your own machine.

The author reports that a 110B model on a 2016 PC with 16GB RAM and SATA storage was predicted to run at 0.2–0.3 tok/s and measured at 0.19 tok/s, while the same “law” yielded 19.3 tok/s for a 30B model on a GTX 1060 6GB. The goal is to make local-model performance more predictable and accessible.

The project is looking for testers and contributors, and the author says the work could become a starting point for further research on token economy and local inference tradeoffs.

Original post →

More from Infra

Infra channel →