Measuring the environmental impact of delivering AI at Google Scale
Cooper Elsworth, Keguo Huang, David Patterson, Ian Schneider, Robert Sedivy, Savannah Goodman, Ben Townsend, Parthasarathy Ranganathan, Jeff Dean, Amin Vahdat, Ben Gomes, James Manyika
cs.AI
2025-08-22
Google’s production telemetry puts a median Gemini Apps text prompt at 0.24 Wh, 0.03 gCO2e, and 0.26 mL of water. Energy fell 33x and carbon 44x in a year, well below most public estimates.
Public estimates for one chatbot prompt span 0.3 to 3 Wh of energy and 0.3 to 45 mL of water. The spread is mostly a measurement-boundary problem. Most studies meter active GPUs, often with batch size pinned at 1. A production fleet also pays for host CPU and DRAM, idle machines kept for latency and failover, and data-center PUE. Google instruments Gemini Apps in production and reports the first full-stack numbers from a large provider.
In the literature, De Vries estimated about 3 Wh from A100s and GPT-3.5. Epoch.AI put a typical GPT-4o prompt near 0.3 Wh on H100s. Altman reported 0.34 Wh and about 0.3 mL of water in 2025 with no boundary. Luccioni et al. measured BLOOM, including GPU, CPU, and DRAM, at about 4 Wh per prompt. The same Llama 3.1 70B can differ by 6x across energy benchmarks.
The functional unit is one serving AI computer: accelerator trays plus a host tray. Energy splits into four buckets: active accelerators (prefill, decode, and on-machine interconnect), active CPU and DRAM, idle machines reserved for spikes and failover, and campus-level PUE overhead.
Telemetry maps job IDs to machines, reads PSU power, then ranks models by energy per prompt and picks the model that serves the 50th-percentile prompt. The distribution is right-skewed, so the paper reports the median rather than the mean. Carbon uses a market-based grid factor (94 gCO2e/kWh in 2024) plus Scope 1+3 for accelerators and hosts. Water uses consumptive WUE of 1.15 L/kWh, counting evaporated water only. A narrower “existing” baseline keeps only active-accelerator energy and subsamples the 10% most efficient data centers. External networking, end-user devices, and training sit outside the boundary.
For a median Gemini Apps text prompt in May 2025:
| Boundary | Energy | Emissions | Water |
| Existing (accelerators, efficient DCs) | 0.10 Wh | 0.02 gCO2e | 0.12 mL |
| Full stack | 0.24 Wh | 0.03 gCO2e | 0.26 mL |
Accelerators are 0.14 Wh (58%), CPU/DRAM 0.06 (25%), idle and overhead 0.02 each. Scaling active-accelerator energy by about 1.72 covers the rest of the stack, versus the 2x overhead some write-ups assume. 0.24 Wh is less than nine seconds of a 100 W television; 0.26 mL is about five drops. That is one to two orders of magnitude below Mistral’s 45 mL and Li et al.’s 10-50 mL.
From May 2024 to May 2025, median energy fell 33x (about 23x from model changes, 1.4x from utilization). Market-based carbon intensity fell another 1.4x and Scope 1+3 fell 36x, for a 44x drop in total carbon per prompt. Fleet PUE is 1.09. Google’s 2023 location-based factor was 366 gCO2e/kWh, 231 of that offset by purchased clean energy, leaving 135 market-based; 2024 was 345 / 251 / 94, about 30% lower intensity. The efficiency stack they credit includes MoE (a small expert subset per token), AQT quantization, speculative decoding, distillation onto Flash and Flash-Lite, and Ironwood TPUs they put at about 30x the energy efficiency of their first public TPU.
Reporting GPU joules alone understates production energy by about 2.4x. The numbers that look small come from batching, speculative decoding, prefill/decode disaggregation, distillation onto Flash-class models, and custom TPUs, not from batch-1 lab runs. Idle machines are about 8% to 10% of the stack, not a fleet sitting half-empty. Cross-model comparisons are meaningless until the boundary is the same.
Nobody outside Google can audit the telemetry. The median drops long traces and poorly utilized models. Market-based accounting credits purchased clean energy; location-based factors are much higher (345 vs 94 gCO2e/kWh in 2024). The 33x energy drop mixes true efficiency with a shift toward distilled smaller models, and the paper does not separate those. Image, video, and extended thinking are out of scope. Training is a separate bill.