OpenAI prepares broader rollout of Ultrafast API offering up to 750 tokens/sec
testingcatalog · x · 2026-09-26
TestingCatalog spotted multiple new references across OpenAI Platform and API docs suggesting a wider rollout of the Ultrafast API mode around DevDay on September 29.
- A hidden speed selector for the Responses API Playground would let developers pick Standard, Fast, or Ultrafast processing.
- OpenAI previously previewed Ultrafast with GPT-5.6 Sol: up to 750 output tokens per second and up to 14× faster inference than Standard.
- Access is currently limited to select customers, and OpenAI confirmed the mode is powered by Cerebras, tied to a 750 MW partnership with staged capacity through 2028.
- For developers the key trade-off is economics: reserving higher-cost, low-latency inference for latency-sensitive workloads.
- It remains unconfirmed whether all GPT-6 family models will support Ultrafast at launch.
Related event: OpenAI Set to Expand Ultrafast API Ahead of DevDay(3 posts)→
More from Infra
- Why rent servers when agents can run your terminal? — StewartalsopIII · 2026-09-26
- MIT paper: scaling law expiring as cost doubles 6x per width jump — DavidLinthicum · 2026-09-26
- Investor thesis: GOES steel and transformer makers may outshine rare earths — basedjensen · 2026-09-26
- Polymorf pushes OMLX inference from 150tps to nearly 200tps within 24 hours of launch — HankYeomans · 2026-09-26
- FlashLoop exploits cross-loop redundancy to speed up Looped Transformers by 1.65x — KyeGomezB · 2026-09-26
- Why this builder quit server racks: fried motherboards and a ~$1,500 housing bill — TheZachMueller · 2026-09-26