Cognition President: Cerebras chips run our models at 950 tokens/s, so fast the team was confused
Sethwinterroth · x · 2026-08-05
Cognition President Russell Kaplan revealed that their deployed Cerebras chips ran their own models at about 950 tokens per second, many times faster than GPUs for the same model size, so fast that the team was confused by the test results. He noted Cerebras occupies a unique point on the price-throughput Pareto curve, and although serving was slightly more expensive, it enabled them to ship a product experience that wouldn't have been possible otherwise.
OpenAI researcher Jeffrey Wang also confirmed that some internal OpenAI models run on Cerebras chips, with inference so fast that tasks finish before he can context-switch, greatly boosting his productivity. He called latency a major bottleneck for useful deployments and expressed excitement about ultra-low-latency inference.
More from Infra
- Together AI's Monthly Token Volume Skyrockets from 30B to 400T — togethercompute · 2026-08-05
- API Key Expiry Leads to Runaway Agent, Costs $300 in Idle Compute — voooooogel · 2026-08-05
- Dev Builds Pixel-Art GPU Cluster Dashboard in 20 Mins Using GLM Agent — Porespellar · 2026-08-05
- DeepSeek-V4-Flash Runs with 256k Context on 4x 4090 GPUs — dangerous_inference · 2026-08-05
- Extropic Founder: Running Probabilistic AI on Deterministic Hardware is a 'Demon Tax' — beffjezos · 2026-08-05
- Samsung Showcases Its 3D NAND and HBM5 Memory Technology — BenBajarin · 2026-08-05