OpenAI Previews Ultrafast Mode for GPT-5.6 Sol at 14X Speed
OpenAI · youtube · 2026-08-14
OpenAI has officially previewed Ultrafast mode, a new service tier designed for the GPT-5.6 Sol model, powered by Cerebras hardware.
- Extreme Speed: The new tier runs up to 14x faster than standard processing, generating up to 750 output tokens per second.
- Workflow Transformation: The massive speed is changing behaviors. A security investigation workflow that previously took 1–2 hours now completes in 10–15 minutes, approaching real-time. Staff also use it to investigate root causes, search systems in parallel, and maintain flow while coding.
- Availability: Ultrafast mode is currently available to a select group of API customers, with access expanding as capacity grows.
Related event: OpenAI and Cerebras Preview GPT-5.6 Sol Ultrafast Mode at 750 Tokens/s(9 posts)→
More from Infra
- What Is the Thermodynamic Limit on Energy Per LLM Token? — prateekj · 2026-08-14
- Prime Flash MoE: Blackwell-Optimized CUDA Kernels for MoE Inference — sloppenheimer · 2026-08-14
- YC-Backed Dipole Labs Uses Optical Switches to Solve AI Compute Idle Time — ycombinator · 2026-08-14
- Cloudflare: The Last Hope of the Open Web Against AI Monopoly? — thedealdirector · 2026-08-14
- TensorSharp vs. llama.cpp: Benchmarking Muse Glimmer 30B Locally — fuzhongkai · 2026-08-14
- Cerebras Teams Up with OpenAI to Massively Accelerate GPT-5.6 Inference — pr337h4m · 2026-08-14