Cerebras CEO says inference speed is reshaping the entire AI chip stack
mattturck · x · 2026-07-24
Why the chip industry is reorganizing around inference speed
Matt Turck’s conversation with Cerebras CEO Andrew Feldman traces the AI chip stack from “what is a wafer?” to the idea that inference latency and tokens-per-second are becoming the key bottleneck.
- The discussion compares GPUs, TPUs, Trainium, and ASICs, and explains why fast inference has turned into a separate race alongside training.
- Feldman argues that the hidden constraints are increasingly about memory and packaging — especially HBM, CoWoS, and 3nm supply.
- The interview also covers prefill vs. decode, why GPUs struggle with decode, and how agents are creating more CPU demand.
- On the business side, Cerebras’ model spans hardware, cloud, and API, with a focus on selling “fast tokens” as a cloud product.
- The conversation touches on OpenAI’s 750MW inference deal, the role of megawatt-scale data centers, and why today’s models may end up being “the worst you ever use.”
Related event: Cerebras CEO: Inference Speed is Reshaping the AI Chip Industry(2 posts)→
More from Companies & People
- AI post-training and evals jobs pay up to $850K, with median offers at $210K–$325K — FinanceYF5 · 2026-09-11
- Only a 4-day window: timeline casts doubt on OpenAI's independent math result claim — gleech · 2026-09-11
- Debate over OpenAI's math breakthrough race: timeline suggests a 4-day turnaround — gleech · 2026-09-11
- 89% of firms use AI, only 6% see significant ROI — the busywork illusion — mikeflache · 2026-09-11
- a16z partner flips to call for nationalizing frontier AI labs, sparking debate — S_OhEigeartaigh · 2026-09-11
- Glean Is Worth $7.2B, but What's Actually Its Moat? — yogthinks · 2026-09-11