Speed Wars: GPT 5.6 Hits 750 Tokens/s on Cerebras

DeepLearningAI · x · 2026-09-02

Top AI companies are prioritizing inference speed as an architectural requirement. OpenAI and Cerebras previewed GPT 5.6 Sol running at 750 tokens per second (11x faster than standard). Google released Gemini 3.7 Flash averaging 330 tokens/s. Nvidia launched Nemotron 3.5 Lightning with dynamic step routing. These improvements aim to reduce context switching and power real-time agentic workflows.

Original post →

More from Infra

Infra channel →