Emad Mostaque says Kimi K3 inference costs could drop 10x to 50x as infra matures
rohanpaul_ai · x · 2026-07-21
Emad Mostaque argues that Kimi K3’s current inference cost is high because the infrastructure around it is still immature, not because of a permanent technical limit. - He expects the cost to fall **10x to 50x** over the next few months as specialized providers optimize **kernels, routing, quantization, batching, memory, and serving**. - He says Kimi K3 currently uses roughly **2x the tokens** of GPT-5.6 for the same task, but that gap should narrow as the ecosystem adapts. - His broader point: U.S.-based inference companies will spend heavily optimizing Chinese open-model weights once they are available, similar to how Fireworks, Modal, and Baseten have built large businesses around inference infrastructure.
Related event: Kimi K3 Reshapes Global AI Pricing, Sparking China-US Compute Cost Debate(9 posts)→
More from Venture
- 430M Organic Impressions: Why LinkedIn Traffic is Worth 10x More Than X — femke_plantinga · 2026-07-21
- Fluidstack raises $830M at $7.5B valuation as Anthropic backs a $50B compute buildout — rohanpaul_ai · 2026-07-21
- Kimi K3 sparks a reset in AI infrastructure value capture, with Bittensor subnets in focus — markjeffrey · 2026-07-21
- Nic Carter argues cheaper AI tokens may break OpenAI and Anthropic’s current business models — markjeffrey · 2026-07-21
- AI growth commitments are disclosed capital spending, not “hidden debt,” says investor — DarinFeinstein · 2026-07-21
- Scott Galloway says Bret Taylor could lead OpenAI after a possible $10B Sierra deal — damianplayer · 2026-07-21