Local inference economics: $60/month power bill for slow speeds
Thin_Pollution8843 · reddit · 2026-08-19
A real-world cost analysis of local inference on a Threadripper 3975WX with 4x V620 GPUs reveals a monthly electricity cost of $55-60 for 6 hours daily use. Despite this, performance with Qwen3.8-27B is sluggish (1-1.3k prefill, 30ts). The author argues that cloud services like OpenRouter or ChatGPT offer better value unless privacy is paramount or solar power is available.
More from Infra
- Local AI apps on a MacBook at zero marginal cost via Pinokio — MLX is alive and well — cocktailpeanut · 2026-08-19
- PA Governor Enacts Nation's Strictest AI Data Center Standards via Executive Order — TinfoilTricorn · 2026-08-19
- Miles v0.1 Open Source RL Framework Launches for LLMs — ying11231 · 2026-08-19
- OpenAI Codex Lead Reveals Tokenizer Inefficiency Can Spike API Bills by 34% — 新智元 · 2026-08-19
- GitHub traffic explosion driven by agents: paid-only hosting isn't the solution — steipete · 2026-08-19
- Input tokens consume 54% of usage in agentic coding, revealing costly 'communication tax' — rseroter · 2026-08-19