Google TPU beats NVIDIA Vera Rubin NVL72 on a third-party model without MTP or disaggregation
YouJiacheng · x · 2026-08-25
A widely shared comparison claims Google's TPU, running in a pure TP setup with no MTP and no prefill/decode disaggregation, still outperformed NVIDIA's Vera Rubin NVL72 on a third-party model using the A0 stepping — and the B0 stepping is reportedly 25% faster. The original poster quipped that NVIDIA GPUs are becoming "HBM wrappers," and the resharing author described the result as "like an atomic bomb going off."
More from Infra
- Leasing a 256GB M5 Ultra Mac Studio now costs barely more than a Claude Max sub — pcuenq · 2026-08-26
- Stagnant US Grid Demand Led to Brain Drain and Lack of Innovation — wandb · 2026-08-26
- OpenAI's Jalapeño inference chip: 1.9x efficiency gain revealed — testingcatalog · 2026-08-26
- Merge Launches Workforce: A Unified Governance Layer for Enterprise AI Access — shensi · 2026-08-26
- Stripe Tests Provisioning API for Agentic App Stores — jeff_weinstein · 2026-08-26
- Monthly API bills may be the wrong metric for LLM budgeting — entelligenceai17 · 2026-08-26