Inference Costs Halved but Quotas Tightened?

mark_k · x · 2026-07-12

This post questions a claim about OpenAI:

> "OpenAI engineers just halved its inference costs."

The author finds it ironic because, simultaneously, many on X are complaining that OpenAI has tightened the usage limits for GPT-5.6, making it feel like "inference is getting more expensive."

Citing The Information, the post notes that OpenAI engineers halving inference costs is their closely guarded "secret weapon," fearing that leaks would allow other labs to optimize their own costs.

Related event: GPT-5.6 Sol Silently Nerfed, OpenAI Confirms Rollback(8 posts)→

Original post →

More from Infra

Infra channel →