128GB local LLM server forges 80M tokens a month on just $6 of electricity

No-Fuel-9202 · reddit · 2026-09-26

A Reddit user details running a 128GB GMKtec X2 as an always-on local LLM server: on balanced performance settings it generates roughly 80 million Qwen3.8 Flash Next tokens per month for about $6 in electricity, while their regular usage is around 30 million tokens monthly. They also used the pi.dev coding agent's autoresearch extension to squeeze performance out of their astrometry app. With optimization work running dry, they're weighing ideas like expanding Karpathy's wiki or RAG-ing their codebase for the idle capacity — and asking the community for better uses.

Original post →

More from coding & agent

coding & agent channel →