OpenAI and Cerebras Preview GPT-5.6 Sol Ultrafast Mode at 750 Tokens/s

机器之心 · wechat · 2026-08-14

OpenAI and AI chip startup Cerebras previewed the 'Ultrafast Mode' for the flagship GPT-5.6 Sol model. Breaking traditional GPU memory bandwidth bottlenecks, the output speed reaches up to 750 tokens/s, a 14x increase over the standard mode without quality degradation.

In the challenging Humanity's Last Exam (HLE) benchmark, the accelerated GPT-5.6 Sol answered all 2,500 PhD-level questions in just 11 hours and 11 minutes, whereas competitor Claude Fable 5 took over 3 days—a 7x speed advantage.

This leap relies on Cerebras' unique wafer-scale architecture (WSE-3). Its 44GB of on-chip SRAM eliminates constant weight fetching from external memory, while pipelining across wafers solves decoding bottlenecks. The ultra-fast inference will massively empower real-time agentic workflows in finance and incident response.

Related event: OpenAI and Cerebras Preview Ultrafast Mode for GPT-5.6 Sol at 750 Tokens/s(20 posts)→

Original post →

More from Fun

Fun channel →