Sarah Hooker: The Inference Era Is Breaking a Decade of GPU Design Assumptions

sarahookr · x · 2026-09-02

Cohere VP of Research Sarah Hooker argues the premises underlying AI hardware are shifting. For a decade, hardware was organized around co-locating as much compute as possible and maximizing matrix-multiply throughput, because pretraining was cost-intensive and inter-GPU data transfer was the most unreliable bottleneck—never let a training run fail. Both assumptions are now changing for two different reasons, reshaping how accelerators and inference infrastructure should be designed.

Related event: Sarah Hooker: Test-Time Compute and Agents Are Forcing an AI Infrastructure Rebuild(9 posts)→

Original post →

More from AGI Musings

AGI Musings channel →