GPT-4-class inference fell 60x in 45 months to $0.33/M tokens, but the floor is rising
rohanpaul_ai · x · 2026-09-17
Rohan Paul breaks down a key inflection in LLM inference economics:
- Cost collapse: GPT-4-class inference dropped from $20 to $0.33 per million tokens in 45 months — a 60x decline far faster than historical PC-compute or bandwidth cost curves
- The floor is rising: DeepSeek raised V4-Pro output prices 2.3x–4.6x
- Kimi K3's sticker-price trap: heavier token usage per task shrinks its apparent discount to only 30% cheaper per completed task than GPT-5.6 Sol — per-task cost, not per-token price, is what matters
He also cites a 91-page Mozilla report: open-weight models are now only 4 months behind the frontier; 8 of OpenRouter's top-10 models by August token volume were open-weight (7 Chinese-built), with DeepSeek the first open model to lead weekly requests — though the economics remain lopsided, with open models handling only 20% of measured usage.
More from Infra
- VC-Attention: training-free low-bit attention hits 1.9x on B200, beating FlashAttention-4 — xiuyu_l · 2026-09-17
- mlx.fast fixes speed-display bug: MLX kernels hit 80.6 tps, nearing 100% speedup milestone — HankYeomans · 2026-09-17
- IBM NorthPole claims 22x inference performance over Nvidia on 12nm process — Site-Staff · 2026-09-17
- Memory shortage hits checkout: Xiaomi raises phone prices 200-1,000 yuan as DRAM stock dips under 10 days — tengyanAI · 2026-09-17
- Burkov Slams OpenRouter Reliability: Fallback Models Fail Together — burkov · 2026-09-17
- Common Crawl puts crawl archives on Hugging Face Storage Bucket, with a getting-started guide — vanstriendaniel · 2026-09-17