Cerebras CTO on the inference frontier: from 4,000 toward 10,000 tokens per second
Latent Space · youtube · 2026-09-03
Latent Space interview with Cerebras co-founder and CTO Sean Lie
- Speed as capability: 100–200 tokens/s may soon feel like "batch mode"; CS4 already pushes inference beyond 4,400 TPS, CS5 targets 10,000 TPS on medium models, and frontier models can run at 5,000 TPS.
- Sold out: Cerebras is effectively sold out of current capacity.
- OpenAI partnership: OpenAI uses ultra-fast inference internally for incident response and critical research; Sean discusses OpenAI's Jalapeño chip, argues AI-first chip design will transform semiconductors, and sees CS5 + Jalapeño combining into a new inference stack.
- Architecture takes: He critiques Groq's SRAM approach at frontier scale, notes models designed around NVIDIA GPUs leave gains on the table for alternatives, and sees hardware-model co-design as the next big win.
- Disaggregation: Inference is splitting into prefill, decode, attention, and expert routing; future data centers may be designed as one giant computer. Memory bandwidth, 3D packaging, power, and cooling are the central bottlenecks.
- US–China: Chinese open models are rising alongside an increasingly independent Chinese AI hardware ecosystem.
More from Infra
- Inference Engineering Is Just a Recipe: vLLM/SGLang, Replicas, Cache-Aware Routing — GabGarrett · 2026-09-03
- Databricks pitches agent-native data infrastructure, Lakebase Postgres at VLDB 2026 — matei_zaharia · 2026-09-03
- Lablup, Maker of GPU Orchestrator Backend.AI, Joins PyTorch Foundation as Silver Member — PyTorch · 2026-09-03
- Cursor cloud agents can now run on your own infrastructure, Mac Minis included — mattyp · 2026-09-03
- VideoDeltaNet open-sources hybrid attention that speeds up MiniMax H3 video generation up to 90x — realmrfakename · 2026-09-03
- Analyst: NVIDIA Could Become Intel Foundry's 'Customer Zero' as a Second Source Beyond TSMC — BenBajarin · 2026-09-03