Google unveils 8th-gen TPU at Hot Chips: two chips per year, split inference and training designs
firstadopter · x · 2026-09-03
- Google presented its 8th-generation TPU at Hot Chips 2026, with TPU tech lead Norman Jouppi and chip architecture senior director Sridhar Lakshmanamurthy presenting.
- Cumulative TPU generations have delivered roughly a million-fold performance increase since the first PCIe inference card.
- Key shift: from one chip per year to two — inference-optimized and training-optimized designs diverge to cover 100k+ chip pre-training runs down to a few-chip inference, distillation, MoE, and agentic models.
- MoE and agentic models demand far more interconnect bandwidth — FLOPs scale easily, interconnects don't — so this generation focused heavily on interconnect scaling.
More from Infra
- Reply reiterating: data centers are good for America's construction workers — saranormous · 2026-09-03
- Should LLM tokens carry green data-center validation labels, like Fair Trade? — jdavid · 2026-09-03
- Commentary: American construction workers want data centers, not just the grey curve — saranormous · 2026-09-03
- Six load forecasters benchmarked on GPU-hours: none beat the last-value baseline — Vegetable-Top-3670 · 2026-09-03
- Visited a 240MW AI data center in Richmond, VA — surprisingly quiet, no high-pitch noise — AndyMasley · 2026-09-03
- Carmack revives rotovators: spinning tethers could slash the cost of space-based data centers — ID_AA_Carmack · 2026-09-03