Running 100 million tokens through GLM 5.2 NVFP4 locally costs about $1
_akhaliq · x · 2026-07-27
Ran 100 million tokens through GLM 5.2 NVFP4 locally for about $1 in electricity, suggesting inference costs can get extremely close to negligible when models are run efficiently on local hardware.
The post shows a cost breakdown of roughly 100.63M tokens for $1.0069, and frames the result as a sign that inference is becoming nearly free.
More from Infra
- Triton backend pushes Falcon3-10B to 97.5 tok/s on an RTX 5070 — OCV_Researcher · 2026-07-27
- Sparrow switches its Standard mode to Ministral 3 14B for local document extraction — andrejusb · 2026-07-27
- Nvidia supplier Wistron opens $700 million Texas plant for GB300 and Vera Rubin systems — Beth_Kindig · 2026-07-27
- BeeLlama.cpp v0.4.1 adds KV-cache precision tails and new quantization modes — Anbeeld · 2026-07-27
- Cloudflare’s AI-training block can also stop Googlebot after September 15 — daluoseo · 2026-07-27
- Self-hosted proxy unifies 424 AI models behind one OpenAI-compatible endpoint — ranadheer535 · 2026-07-27