SemiAnalysis says better memory and storage can beat a faster GPU in modern inference

rwang07 · x · 2026-07-27

Dylan Patel of SemiAnalysis argues that inference is no longer won by buying the newest GPU alone: a weaker GPU with better storage and memory can outperform a faster chip in some workloads.

The thread also says SemiAnalysis now runs more than $80 million of compute across Nvidia GPUs, AMD GPUs, Google TPUs, and Amazon Trainium, with automated daily benchmarks on the latest inference engines, drivers, PyTorch versions, and major Chinese models such as GLM, Zhipu, Moonshot, Kimi, and Alibaba models.

A second claim in the clip is that Nvidia controls a network of 100+ neo-clouds by allocating chips selectively and backstopping loans, using that ecosystem to amplify its message and make alternative hardware harder to deploy.

Overall, the post is about how memory, storage, and distribution power now matter as much as raw GPU speed in inference economics.

Original post →

More from Infra

Infra channel →