Neoclouds add inference, inference firms hoard GPUs: it's all about controlling token flow

AccBalanced · x · 2026-09-28

jessiedong observes vertical integration across the AI stack: neoclouds are adding inference, inference companies are reserving or buying GPUs, and even routers like OpenRouter and Vercel AI Gateway need guaranteed GPU capacity.

The surface motive is money, but the deeper logic, she argues, is controlling where tokens go — which requires controlling GPU supply:

As compute gets scarcer, this integration race should accelerate.

Original post →

More from Infra

Infra channel →