TPU Beats B200 by Up to 50% Per Dollar in First Third-Party Ironwood Inference Benchmark by SemiAnalysis
新智元 · wechat · 2026-09-10
SemiAnalysis published the first third-party inference benchmark of Google's 7th-gen TPU Ironwood: under identical workloads (same open-source model, FP8 vs FP8, 8k-in/1k-out), TPU delivers up to 50% better performance per dollar than NVIDIA B200 and nearly 96% vs B300.
Key numbers
- At 100 token/s single-user throughput, TPU is 19% cheaper than B200 and 34% cheaper than B300
- Cost per million tokens: TPU $0.181, B200 $0.222, B300 $0.276
- On Google's internal $1.03/chip-hour TCO basis, the advantage widens to 77%–130%
Engineering: Ironwood uses dual compute dies, 6x HBM vs Trillium, and native FP8; doubling KV-cache pages cut median time-to-first-token by 95%; splitting read/compute blocks lifted decode throughput from 64.9k to 96.3k token/s; offloading ReduceScatter and MoE routing to SparseCore gained up to 26.1% throughput.
Ecosystem: TorchTPU bypasses the TorchAX translation wall via PyTorch's PrivateUse1 backend, letting vLLM/SGLang reuse schedulers and API layers directly. The piece also dismantles Jensen Huang's April claims that "TPU won't benchmark" and "Anthropic is an outlier" — Anthropic has committed to over 1 million TPUs (400k purchased, 600k leased) and will become Google's largest TPU customer by 2029.
More from Infra
- M.2-Oculink eGPU Link Silently Downgrades to PCIe Gen1 — Here's How to Check — El_90 · 2026-09-11
- DeepSeek launches V4.1-Flash with 1M-token context and 4x smaller KV-cache — matlabulous · 2026-09-11
- What Can You Still Run on 8GB VRAM? User Asks for Small Models With Tool Use — riceinmybelly · 2026-09-11
- Spain's hourly 80% renewable matching rules clash as France fast-tracks 700MW sites, UK cuts grid queues — eherrerosj · 2026-09-11
- AI could add 0.3-0.4 points to Europe's productivity growth, but the EU holds under 5% of global compute — rohanpaul_ai · 2026-09-11
- Qualcomm's next-gen Hexagon NPU runs 30B MoE models with 32K context on-device — lee_stott · 2026-09-11