Neoclouds add inference, inference firms hoard GPUs: it's all about controlling token flow
AccBalanced · x · 2026-09-28
jessiedong observes vertical integration across the AI stack: neoclouds are adding inference, inference companies are reserving or buying GPUs, and even routers like OpenRouter and Vercel AI Gateway need guaranteed GPU capacity.
The surface motive is money, but the deeper logic, she argues, is controlling where tokens go — which requires controlling GPU supply:
- Neoclouds add inference to maximize GPU utilization
- Inference companies own compute so AWS or a neocloud can't block them from serving requests
- Routers need providers with available GPUs or requests have nowhere to go
As compute gets scarcer, this integration race should accelerate.
More from Infra
- FailureAtlas: most severe LLM gateway failures return HTTP 200 and silently corrupt your app — its_vayishu · 2026-09-28
- Used RTX 3090 prices creep toward $1,500 on eBay amid GPU shortage — sleight42 · 2026-09-28
- gufo inference doubles prefill speed vs llama.cpp forks for Qwen 3.8 Flash Next on Strix Halo — fallingdowndizzyvr · 2026-09-28
- Scraping p50 stabilized at 2s: keep your app and databases colocated — DanielLockyer · 2026-09-28
- PKU open-sources RayOrch, lineage-aware data-prep engine with up to 15.14x speedup — PekingUniversity · 2026-09-28
- Why credit quality matters most in compute: clusters are opaque, people are trackable — AccBalanced · 2026-09-28