Engy posts live inference prices as Qwen3.6 undercuts GLM-5.2 on cached input
markjeffrey · x · 2026-07-22
Engy published live per-token pricing for its inference service, including cached-input rates that are far below headline pricing.
The screenshot shows glm-5.2 at $0.68 per million input tokens and $1.50 per million output tokens, while qwen3.6-35b-a3b is priced at $0.045 input / $0.30 output, with cached input at $0.015. The page also notes that prompt-cache hits are billed automatically at the cached rate and that agentic workloads often achieve 90%+ cache on repeated prefixes, making effective costs much lower than the listed rates.
More from Infra
- Mark Cuban says cheaper AI could leave data centers empty enough for pickleball — AIFlow_ML · 2026-07-22
- Dongfang Computing Core’s DF1000 chip claims 520 TFLOPS BF16 and 6.4 TB/s bandwidth — teortaxesTex · 2026-07-22
- QuixiAI shows the same runtime spanning CUDA, Metal, ROCm, XPU, Gaudi and CPU — QuixiAI · 2026-07-22
- Meta infra is accused of wasting silicon on local wins that cost billions — dylan522p · 2026-07-22
- AMD teases an AI event with a “Build What’s Next” banner — xiaosun86 · 2026-07-22
- Azure Architecture Diagram Builder adds MCP support for agent-driven Bicep workflows — davemccollough · 2026-07-22