Running 100 million tokens through GLM 5.2 NVFP4 locally costs about $1

_akhaliq · x · 2026-07-27

Ran 100 million tokens through GLM 5.2 NVFP4 locally for about $1 in electricity, suggesting inference costs can get extremely close to negligible when models are run efficiently on local hardware.

The post shows a cost breakdown of roughly 100.63M tokens for $1.0069, and frames the result as a sign that inference is becoming nearly free.

Original post →

More from Infra

Infra channel →