AI is cheaper to start, but 23.4x more expensive to run continuously
0xJeff · x · 2026-07-22
The cheap-to-start, expensive-to-run problem is getting worse
A repost argues that AI has become far cheaper to start using, but much more expensive to keep running in production:
- Inference cost to get started has fallen from $18.40 to $6.07 per 1M tokens in a year.
- But the cost of continuously running agentic workflows is now 23.4x higher.
- As users adopt AI more deeply, a single engineer may go from using Claude for code review to using AI for coding, travel booking, personal finance, and more.
- The poster also says DeepSeek doubled prices during peak China hours, and that Kimi was difficult to subscribe to because capacity was sold out.
The attached chart from OpenRouter shows rapid growth in model usage, with MiMo-V2.5, DeepSeek V4 Flash, Hy3, MiniMax M3, and GLM 5.2 among the top used models.
More from Infra
- 12 KV Cache Reduction Techniques Every AI Engineer Should Understand, Explained — blaizedsouza · 2026-09-11
- The shadow GPU capacity market is formalizing, with Meta selling excess compute to outside buyers — DavidLinthicum · 2026-09-11
- Engram's random reads don't suit SSDs; CPU-memory over NVLink could serve all 72 GPUs — bookwormengr · 2026-09-11
- 80% of the DIY LLM inference hype posters have already quit — it's brutally hard systems work — abhijithneil · 2026-09-11
- Hugging Face's Ultra Scale Playbook: a free book on training LLMs on GPU clusters — mdancho84 · 2026-09-11
- Is inference latency becoming the biggest bottleneck for production AI agents? — Euphoric_Sea632 · 2026-09-11