OpenAI Slashes GPT-5.6 Prices by 80%, Inference Cost Drops 2000x Annually
Latent Space · rss · 2026-07-31
OpenAI announced significant price cuts for GPT-5.6: Luna down 80%, Terra down 20%, and a new Sol Fast mode with 2.5x lower latency at 2x price. System-level optimizations like self-optimizing kernels, speculative decoding, and KV caching reduced serving costs. GPT-5.4 flagship intelligence now costs 1/13th of its price four months ago, an annualized 2000x drop. ARC-AGI-3 discussions emphasize evaluating the full agent system, not just model weights.
More from Infra
- metal-graph 0.1.0: Fast Graph Analytics on Apple Silicon via Metal — HankYeomans · 2026-07-31
- Deep Dive into DeepSpeedEngine: Architecting a God-Object for Complex Training — Mahmoud_Zalt · 2026-07-31
- How to Run Qwen on 3x 2080Ti and 128GB RAM? Local Deployment Help — AccountGotLocked69 · 2026-07-31
- The Cost of 'Good Enough' Data: Why Modern Architectures Fail at Scale — craigmullins · 2026-07-31
- Big Tech AI spending tops $1 trillion, FT reports — gaganghotra_ · 2026-07-31
- Revisiting Lossy Verification in Speculative Decoding: Mechanisms and Failure Modes — Tianyu Wang · 2026-07-31