SemiAnalysis: TPU v7 Ironwood beats Blackwell Ultra by 50% perf per dollar on inference
Sentdex · x · 2026-09-08
SemiAnalysis reports that on apples-to-apples InferenceX benchmarks, Google's TPU v7 Ironwood delivers 50% better performance per dollar than Nvidia's Blackwell Ultra when using the new TorchTPU inference stack.
TorchTPU brings native PyTorch support to TPUs, and with Google open-sourcing a set of Pallas inference kernels, the firm argues TPU externalization is moving full steam ahead. Sentdex quips that SemiAnalysis is like Ed Zitron, but just for Nvidia.
More from Infra
- Hyperscaler backlog hits $1.7T as Citi conference flags shift to agentic inference — sanjaykalra · 2026-09-08
- AI data center interconnect chip startup Celero raises $275M at $3B+ valuation — dinabass · 2026-09-08
- vLLM's Speculators v0.8.0 ships unified CLI, PyPI Mooncake connectors, fused Triton loss kernel — vllm_project · 2026-09-08
- Dev buys a Mac mini just to run Codex 24/7 across his entire workflow — _AustinCalvert_ · 2026-09-08
- INT21's agent-generated Qwen3.8 trainer hits 11.5x PyTorch FSDP2 throughput on 8 B200s — bingxu_ · 2026-09-08
- Benchmarking Gemma under 16-128 concurrent users: classification vs generation loads — rseroter · 2026-09-08