TPU Beats B200 by Up to 50% Per Dollar in First Third-Party Ironwood Inference Benchmark by SemiAnalysis

新智元 · wechat · 2026-09-10

SemiAnalysis published the first third-party inference benchmark of Google's 7th-gen TPU Ironwood: under identical workloads (same open-source model, FP8 vs FP8, 8k-in/1k-out), TPU delivers up to 50% better performance per dollar than NVIDIA B200 and nearly 96% vs B300.

Key numbers

Engineering: Ironwood uses dual compute dies, 6x HBM vs Trillium, and native FP8; doubling KV-cache pages cut median time-to-first-token by 95%; splitting read/compute blocks lifted decode throughput from 64.9k to 96.3k token/s; offloading ReduceScatter and MoE routing to SparseCore gained up to 26.1% throughput.

Ecosystem: TorchTPU bypasses the TorchAX translation wall via PyTorch's PrivateUse1 backend, letting vLLM/SGLang reuse schedulers and API layers directly. The piece also dismantles Jensen Huang's April claims that "TPU won't benchmark" and "Anthropic is an outlier" — Anthropic has committed to over 1 million TPUs (400k purchased, 600k leased) and will become Google's largest TPU customer by 2029.

Original post →

More from Infra

Infra channel →