Open-source inference cloud operator: compute demand will be desperately short if workloads shift to open models
gharik · x · 2026-10-08
Jeffrey Huang says he's deeply involved in a decently sized inference cloud that runs only open-source models, including Chinese ones. His take: if compute shifts away from Anthropic and OpenAI toward open models, inference compute will be desperately needed — "we already do".
It's a counterpoint to the narrative that open models reduce compute demand: open-weight inference still burns massive compute, likely less efficiently.
More from Infra
- Tencent's STEPQuant: 6-bit recurrent states match FP32 with 68.7% less memory — _akhaliq · 2026-10-09
- Bain projects 38.6M GPU and custom silicon shipments by 2030, 10x 2023's 3.9M — Beth_Kindig · 2026-10-09
- AWS reference architecture: multi-team GPU cluster sharing on SageMaker HyperPod — AWS ML Blog · 2026-10-09
- Mistral slammed for training open models on datacenters powered ~70% by coal — wavefnx · 2026-10-09
- Why do we resend the whole conversation every turn? Server-side KV slots proposal sparks debate — Vasili_Sk · 2026-10-09
- NVIDIA's NeMo-DCR cuts 1T-model weight sync from 87.5 min to 150s, 12-40x faster checkpoint transfer — dair_ai · 2026-10-09