Cognition President: Cerebras chips run our models at 950 tokens/s, so fast the team was confused

Sethwinterroth · x · 2026-08-05

Cognition President Russell Kaplan revealed that their deployed Cerebras chips ran their own models at about 950 tokens per second, many times faster than GPUs for the same model size, so fast that the team was confused by the test results. He noted Cerebras occupies a unique point on the price-throughput Pareto curve, and although serving was slightly more expensive, it enabled them to ship a product experience that wouldn't have been possible otherwise.

OpenAI researcher Jeffrey Wang also confirmed that some internal OpenAI models run on Cerebras chips, with inference so fast that tasks finish before he can context-switch, greatly boosting his productivity. He called latency a major bottleneck for useful deployments and expressed excitement about ultra-low-latency inference.

Original post →

More from Infra

Infra channel →