128GB local LLM server forges 80M tokens a month on just $6 of electricity
No-Fuel-9202 · reddit · 2026-09-26
A Reddit user details running a 128GB GMKtec X2 as an always-on local LLM server: on balanced performance settings it generates roughly 80 million Qwen3.8 Flash Next tokens per month for about $6 in electricity, while their regular usage is around 30 million tokens monthly. They also used the pi.dev coding agent's autoresearch extension to squeeze performance out of their astrometry app. With optimization work running dry, they're weighing ideas like expanding Karpathy's wiki or RAG-ing their codebase for the idle capacity — and asking the community for better uses.
More from coding & agent
- No video model needed: Claude Opus 5.5 makes a $2 startup promo video in 1 minute — FinanceYF5 · 2026-09-26
- Anthropic made claude.ai 3x faster in 2 weeks; dev turns the playbook into a speed skill — AlchainHust · 2026-09-26
- 4D docs MCP server brings cached command documentation to AI coding assistants — modelcontextprotocol · 2026-09-26
- AI agents ranked by build difficulty: computer-use hardest, email agents easiest — sahilypatel · 2026-09-26
- BrainAPI Open-Sources a Knowledge Graph–Powered Memory Layer for AI Agents — adnan_hashmi · 2026-09-26
- Your agent didn't break — your delegation model did: the agent-permission problem — Fantastic-Sleep-3352 · 2026-09-26