No one matches its inference economics; commentator suggests 10-20% cut from infra providers

zephyr_z9 · x · 2026-09-10

A discussion around planning a large-scale deployment with 2,000 GPUs plus a storage cluster, with the original poster noting the company's inference architecture is too complex for most infra providers to replicate.

The commenter argues the company should do deployments itself and could even take a 10%-20% cut from infrastructure providers. The broader point: nobody has figured out how to beat or even match its inference economics, so self-deployment is the logical next step.

Original post →

More from Infra

Infra channel →