OpenAI credits caching and inference gains for GPT-6's 50% price cut
OpenAIDevs · x · 2026-09-23
OpenAI Devs explained GPT-6 Sol and Luna's lower prices come from caching and inference improvements, passing savings on: more intelligence at half the previous generation's API prices, making complex automation practical at scale.
Related event: OpenAI Launches GPT-6 Sol and Luna with 50% API Price Cut(6 posts)→
More from Infra
- Dev laments agents built around KV caches, wants inference-first chips — dbreunig · 2026-09-23
- GE Vernova seen hitting $200B backlog by early 2027 as turbine demand outruns guidance — BenBajarin · 2026-09-23
- New LLM papers: recursive language models generalize out of domain; XMerge depth compression — burny_tech · 2026-09-23
- DeepSeek report praised: absurdly tiny KV cache, dense infra design — stochasticchasm · 2026-09-23
- Running Qwen 27B locally on RTX 4090: beats pre-2025 coding models, RAM is the wall — julianharris · 2026-09-23
- FP8 Tuning Cuts 42.9ms Per Step: Custom SGLang Kernels Boost Inference 126% — HankYeomans · 2026-09-23