Sarah Hooker: The Inference Era Is Breaking a Decade of GPU Design Assumptions
sarahookr · x · 2026-09-02
Cohere VP of Research Sarah Hooker argues the premises underlying AI hardware are shifting. For a decade, hardware was organized around co-locating as much compute as possible and maximizing matrix-multiply throughput, because pretraining was cost-intensive and inter-GPU data transfer was the most unreliable bottleneck—never let a training run fail. Both assumptions are now changing for two different reasons, reshaping how accelerators and inference infrastructure should be designed.
More from AGI Musings
- The infiltrator's burden: paranoid AI agents from the HF attack give defenders an asymmetric edge — robleclerc · 2026-09-03
- '90% of work is too vague to even be wrong': researchers slam rigor turning into entertainment — RexDouglass · 2026-09-03
- Hospitals Are Using AI to Raise Prices, Argues Antitrust Commentator Matt Stoller — GaryMarcus · 2026-09-03
- Cooperative AI Summer School 2026 releases talks, gathered 45 early-career researchers — xuanalogue · 2026-09-03
- AI diffusion debate: digital tech now spreads with or without human adoption — soleio · 2026-09-03
- AI agents break the internet's three-layer transaction fraud validation chain — arampell · 2026-09-03