OpenAI Slashes GPT-5.6 Prices by 80%, Inference Cost Drops 2000x Annually

Latent Space · rss · 2026-07-31

OpenAI announced significant price cuts for GPT-5.6: Luna down 80%, Terra down 20%, and a new Sol Fast mode with 2.5x lower latency at 2x price. System-level optimizations like self-optimizing kernels, speculative decoding, and KV caching reduced serving costs. GPT-5.4 flagship intelligence now costs 1/13th of its price four months ago, an annualized 2000x drop. ARC-AGI-3 discussions emphasize evaluating the full agent system, not just model weights.

Original post →

More from Infra

Infra channel →