Inference businesses may be charging 7× to 15× more than renting a GPU

JoshPurtell · x · 2026-07-27

AI inference businesses are still posting 90%–99.9% gross token margins

The post argues that many AI companies are underpricing inference by a wide margin. In the quoted example, renting a GPU for a Qwen3.6 27B bulk inference job was reportedly 7×–15× cheaper than using open inference providers.

Original post →

More from Infra

Infra channel →