Emad Mostaque says Kimi K3 inference costs could drop 10x to 50x as infra matures
rohanpaul_ai · x · 2026-07-21
Emad Mostaque argues that Kimi K3’s current inference cost is high because the infrastructure around it is still immature, not because of a permanent technical limit.
- He expects the cost to fall 10x to 50x over the next few months as specialized providers optimize kernels, routing, quantization, batching, memory, and serving.
- He says Kimi K3 currently uses roughly 2x the tokens of GPT-5.6 for the same task, but that gap should narrow as the ecosystem adapts.
- His broader point: U.S.-based inference companies will spend heavily optimizing Chinese open-model weights once they are available, similar to how Fireworks, Modal, and Baseten have built large businesses around inference infrastructure.
Related event: Kimi K3 Reshapes Global AI Pricing, Sparking China-US Compute Cost Debate(9 posts)→
More from Venture
- Investor argues Palantir-Nvidia partnership should slash Anthropic's IPO valuation — pdamodaran · 2026-09-11
- Moonshot's annualized revenue jumped from $300M to $1B in two months after Kimi K3 — Hesamation · 2026-09-11
- Mid-market companies' AI SEO bottleneck is ops execution, not strategy, says SEO practitioner — gaganghotra_ · 2026-09-11
- A YouTuber with 1.5M followers paid this indie maker for a consulting call — tibo_maker · 2026-09-11
- 71% of people have never used generative AI — the bubble argument for microsaas — iamaliveix · 2026-09-11
- Glean grew from $100M to $300M ARR in roughly fifteen months — yogthinks · 2026-09-11