Inference Costs Halved but Quotas Tightened?
mark_k · x · 2026-07-12
This post questions a claim about OpenAI:
> "OpenAI engineers just halved its inference costs."
The author finds it ironic because, simultaneously, many on X are complaining that OpenAI has tightened the usage limits for GPT-5.6, making it feel like "inference is getting more expensive."
Citing The Information, the post notes that OpenAI engineers halving inference costs is their closely guarded "secret weapon," fearing that leaks would allow other labs to optimize their own costs.
Related event: GPT-5.6 Sol Silently Nerfed, OpenAI Confirms Rollback(8 posts)→
More from Infra
- Tabul AI launches Metal TreeSHAP to speed up Shapley values on Apple silicon — Scobleizer · 2026-07-22
- DeepSeek-V4-Flash tops out at 770 tok/s on one B300 in a vLLM batch test — Moreh · 2026-07-22
- NVIDIA starts shipping 102.4 Tbps Spectrum-6 switches for Vera Rubin AI factories — nvidia · 2026-07-22
- Apple publishes SOC 3 audit reports for Private Cloud Compute — throwfaraway4 · 2026-07-22
- Reddit GPU renters say existing platforms only give you two of three: code, recovery, fair billing — legendpizzasenpai · 2026-07-22
- The Sandboxing Manifesto: Secure Execution Environments for Agents — spirosoik · 2026-07-22