Google TPU beats NVIDIA Vera Rubin NVL72 on a third-party model without MTP or disaggregation

YouJiacheng · x · 2026-08-25

A widely shared comparison claims Google's TPU, running in a pure TP setup with no MTP and no prefill/decode disaggregation, still outperformed NVIDIA's Vera Rubin NVL72 on a third-party model using the A0 stepping — and the B0 stepping is reportedly 25% faster. The original poster quipped that NVIDIA GPUs are becoming "HBM wrappers," and the resharing author described the result as "like an atomic bomb going off."

Original post →

More from Infra

Infra channel →