Rent GPUs, self-serve 90% of tokens: dev slams unbounded LLM API pricing

TheZachMueller · x · 2026-10-10

A developer (@pranavsf) hinted their company faces budget strain as LLM spend outgrows its business model, especially with RL demand rising. Zach Mueller argues more companies will realize it's cheaper to rent GPUs and serve 90% of token spend internally at a fixed rate, leaving only the final 10% to API providers whose uncapped pricing feels exploitative.

Original post →

More from Infra

Infra channel →