Google's AI infra chief: at 100K accelerators, FLOPS is a vanity metric — goodput is what matters
Training Data (Sequoia) · rss · 2026-10-06
Sequoia's Training Data podcast hosts Amin Vahdat, Google's Chief Technologist for AI Infrastructure, on the physics and economics of frontier AI.
- Goodput, not FLOPS: at 100,000-accelerator scale, something fails multiple times an hour, so he argues FLOPS is a vanity metric; the real metric is goodput — useful work delivered through real failures.
- The TPU bet: from a contrarian call in 2013 to the first split into 8i (inference) and 8t (training), plus why TPU core primitives haven't changed since v1.
- Co-design with DeepMind: the teams sit in the same rooms and can intercept chip architectures mid-flight before tape-out.
- Agents reshape the data center: long-horizon agents are sending CPU and storage demand through the roof alongside accelerators.
- Networking and power: optical circuit switches reroute light to a spare rack in milliseconds; power is the binding constraint — Google would rather wait on a utility than build its own gigawatt.
- Also covered: token capacity doubling every six months, orbital data centers, and the multi-megawatt rack of 2036.
More from Infra
- Drax datacentre would burn 4.9m tonnes of wood a year, emissions near double Gatwick flights — nordicinst · 2026-10-06
- Singapore data center operator DayOne files for US IPO after H1 revenue tripled to $512M — zephyr_z9 · 2026-10-06
- Strata claims 6GB VRAM can match RTX 5090-level inference, with full Qwen4 support planned — lxfater · 2026-10-06
- ComfyUI benchmark: Flux 1 dev FP8 tested across 13 GPUs — Ok_Contribution8157 · 2026-10-06
- ODS offers one-click local AI: auto hardware detection, model setup, agents and plugins — Teknium · 2026-10-06
- PlanetScale Engineer Explains Kubernetes Feedback Loops by Running Postgres by Hand — bibryam · 2026-10-06