Cerebras CEO says inference speed is reshaping the entire AI chip stack
mattturck · x · 2026-07-24
Why the chip industry is reorganizing around inference speed
Matt Turck’s conversation with Cerebras CEO Andrew Feldman traces the AI chip stack from “what is a wafer?” to the idea that inference latency and tokens-per-second are becoming the key bottleneck.
- The discussion compares GPUs, TPUs, Trainium, and ASICs, and explains why fast inference has turned into a separate race alongside training.
- Feldman argues that the hidden constraints are increasingly about memory and packaging — especially HBM, CoWoS, and 3nm supply.
- The interview also covers prefill vs. decode, why GPUs struggle with decode, and how agents are creating more CPU demand.
- On the business side, Cerebras’ model spans hardware, cloud, and API, with a focus on selling “fast tokens” as a cloud product.
- The conversation touches on OpenAI’s 750MW inference deal, the role of megawatt-scale data centers, and why today’s models may end up being “the worst you ever use.”
More from Companies & People
- A post predicts OpenAI will preview an AI device in 2026 and ship it in 2027 — imjustnewatai · 2026-07-24
- NeurIPS author says this year’s reviews were notably more helpful — chhaviyadav_ · 2026-07-24
- Anthropic says Claude helped bootstrap an AMD Instinct MI355 rack on ROCm — BenBajarin · 2026-07-24
- AMD pitches 256-core Venice as a CPU for agent sandboxes and AI host nodes — BenBajarin · 2026-07-24
- Leaked Transcript Reveals DeepSeek CEO's Views on US-China AI Race — LuizaJarovsky · 2026-07-24
- AMD unveils Venice server CPUs with up to 256 cores and OpenAI tie-in — ryanshrout · 2026-07-24