Inference is turning GPU compute into a tradable commodity
ArtificialAnlys · x · 2026-09-09
jessiedong argues that inference is already making compute tradable: inference providers are "partly compute traders" — they reserve compute, run models on it, and sell output by the token, profiting by buying compute cheaply and squeezing more tokens per GPU while taking idle-capacity risk.
Providers turn disparate hardware into an easily comparable commodity (the same model's tokens at set price and speed). The market already shows a division of labor:
- GPU capacity → tokens: Together, Fireworks, DeepInfra, Baseten
- Own chips → tokens: Groq, Cerebras, SambaNova
- Price/speed comparison: Artificial Analysis
- Cross-provider comparison and routing: OpenRouter and others
If you're bearish on trading GPUs, look at inference instead.
More from Venture
- Legal AI startup Legora passes $100M ARR with 10x year-over-year growth — ycombinator · 2026-09-09
- Runway hits $200M ARR, targeting $1B next — tlakomy · 2026-09-09
- OpenAI CFO says ARR up 20% in a month, ads fastest to $1B, compute push continues — Kr00ney · 2026-09-09
- levelsio makes his decade-old community nearly free: $1 signup, here's why — PratikKadam_ · 2026-09-09
- AI search visibility is about what third parties say about you, not your site — nikvassev · 2026-09-09
- Ramp builds semantic layer linking AI agent spend to actual business ROI — floguo · 2026-09-09