OpenAI Previews GPT-5.6 Sol Ultrafast Mode: 14x Speed Boost Powered by Cerebras
Justgototheeffinmoon · reddit · 2026-08-16
OpenAI has launched a limited preview of an 'Ultrafast' mode for GPT-5.6 Sol, powered by its partnership with Cerebras. The mode delivers up to a 14x speed increase, reaching output speeds of 750 tokens per second. Preview customers, including Jane Street, are testing it in production for coding, financial research, and support use cases. OpenAI is also using it internally for incident response and rapid research iteration. While the figures are impressive, pricing, general availability dates, and performance under real concurrency remain undisclosed. This development signals that specialty silicon is becoming a core part of OpenAI's product stack, potentially reshaping latency budgets for interactive AI products.
More from Infra
- MusCoRe protocol cuts agent history tokens by 71.9%, targeting edge inference on Pi 5 — Parallel_News · 2026-08-16
- Developer investigates undocumented B200 instructions for potential speedups — SkyLi0n · 2026-08-16
- Beating cuBLAS by 4.7%: NVFP4 Kernels Hand-Built on GB300, 100% Claude-Generated — pranjalssh · 2026-08-16
- Qwen 3.8 runs 151% faster on Apple Silicon via community challenge — gajesh · 2026-08-16
- a16z: NeoClouds Repurpose Crypto Infrastructure for AI Compute Boom — a16z · 2026-08-16
- Colibri: Run 2.8T-parameter MoE models on a 25GB laptop with pure C inference engine — alex_verem · 2026-08-16