Engy posts live inference prices as Qwen3.6 undercuts GLM-5.2 on cached input

markjeffrey · x · 2026-07-22

Engy published live per-token pricing for its inference service, including cached-input rates that are far below headline pricing.

The screenshot shows glm-5.2 at $0.68 per million input tokens and $1.50 per million output tokens, while qwen3.6-35b-a3b is priced at $0.045 input / $0.30 output, with cached input at $0.015. The page also notes that prompt-cache hits are billed automatically at the cached rate and that agentic workloads often achieve 90%+ cache on repeated prefixes, making effective costs much lower than the listed rates.

Original post →

More from Infra

Infra channel →