Inference businesses may be charging 7× to 15× more than renting a GPU
JoshPurtell · x · 2026-07-27
AI inference businesses are still posting 90%–99.9% gross token margins
The post argues that many AI companies are underpricing inference by a wide margin. In the quoted example, renting a GPU for a Qwen3.6 27B bulk inference job was reportedly 7×–15× cheaper than using open inference providers.
- The author says real-world gross token margins of 90%, 95%, 99.5%, and even 99.9% are out there.
- The cited comparison suggests open providers can be dramatically more expensive than self-rented GPUs for batch workloads.
- The implication is that many users may be getting “hosed” on inference pricing, especially when workloads are large and predictable.
Related event: AI Inference Services Cost Up to 15x More Than Renting GPUs(4 posts)→
More from Infra
- Together AI adds canary rollouts for zero-downtime model upgrades on dedicated inference — togethercompute · 2026-09-23
- Dedicated Hardware for Running AI Agents at Scale Arrives — cyrilzakka · 2026-09-23
- Ternary Bonsai 2 27B: 5.9GB weights retain ~95% of full-precision reasoning — cephaloform · 2026-09-23
- Qwen 27B runs 24hr unattended on one RTX5090, builds full Postgres-SpringBoot-React spreadsheet app — anglepoiselife · 2026-09-23
- OpenRoboto Shift launches: decentralized egocentric video data network for robot brains — markjeffrey · 2026-09-23
- Engineer describes designing digital circuits that recycle most of their energy — MikePFrank · 2026-09-23