Inference businesses may be charging 7× to 15× more than renting a GPU
JoshPurtell · x · 2026-07-27
AI inference businesses are still posting 90%–99.9% gross token margins
The post argues that many AI companies are underpricing inference by a wide margin. In the quoted example, renting a GPU for a Qwen3.6 27B bulk inference job was reportedly 7×–15× cheaper than using open inference providers.
- The author says real-world gross token margins of 90%, 95%, 99.5%, and even 99.9% are out there.
- The cited comparison suggests open providers can be dramatically more expensive than self-rented GPUs for batch workloads.
- The implication is that many users may be getting “hosed” on inference pricing, especially when workloads are large and predictable.
More from Infra
- NVIDIA says AdamW hits a scale ceiling as SOAP and Muon beat it on trillion-token runs — omarsar0 · 2026-07-27
- Local LLM builders ask whether RTX Ada workstation cards are worth tracking — egudegi · 2026-07-27
- Guide breaks down how to cut Microsoft Fabric capacity costs on Azure — adnan_hashmi · 2026-07-27
- Introductory guide explains the `sempy.fabric` package for Microsoft Fabric — adnan_hashmi · 2026-07-27
- Voice AI’s real call cost is more than minutes: STT, TTS, SIP and retries — decant338 · 2026-07-27
- Micron and Meta paper says Spark can slow down 38× when shuffle spills to SSD — dr_alphalyrae · 2026-07-27