No one matches its inference economics; commentator suggests 10-20% cut from infra providers
zephyr_z9 · x · 2026-09-10
A discussion around planning a large-scale deployment with 2,000 GPUs plus a storage cluster, with the original poster noting the company's inference architecture is too complex for most infra providers to replicate.
The commenter argues the company should do deployments itself and could even take a 10%-20% cut from infrastructure providers. The broader point: nobody has figured out how to beat or even match its inference economics, so self-deployment is the logical next step.
More from Infra
- Google Cloud user hit with an $82k bill within 5 hours — Patient_Election2179 · 2026-09-10
- TRL ships 1M-token long-context training guide, trains Qwen3-8B on one 8-GPU node — QGallouedec · 2026-09-10
- Dual RTX Pro 6000 + Threadripper 9955W local LLM build — sanity check requested — No_Run8812 · 2026-09-10
- Screenshot surfaces rare admission of 72-hour KV cache limits in V4-era architecture — zephyr_z9 · 2026-09-10
- DeepSeek cut per-token KV cache size by 54x in nine months — zephyr_z9 · 2026-09-10
- vLLM Ships Full Support for DeepSeek-V4.1-Flash's New Architecture — vllm_project · 2026-09-10