GPT-5.6 Sol Leak: 750 tps

新智元 · wechat · 2026-07-09

Xinzhiyuan reports that OpenAI plans to launch the frontier model GPT-5.6 Sol on Cerebras custom hardware this month, claiming inference speed up to 750 tokens/s, initially available to a few specific customers.

The article speculates this may be a massive multi-wafer deployment, with OpenAI possibly restructuring the model architecture around the hardware to reduce KV cache pressure and enhance real-time agent task experience.

Related event: Reports Say Cerebras Could Push GPT-5.6 to 750 TPS(4 posts)→

Original post →

More from Infra

Infra channel →