Emad Mostaque says Kimi K3 inference costs could drop 10x to 50x as infra matures

rohanpaul_ai · x · 2026-07-21

Emad Mostaque argues that Kimi K3’s current inference cost is high because the infrastructure around it is still immature, not because of a permanent technical limit. - He expects the cost to fall **10x to 50x** over the next few months as specialized providers optimize **kernels, routing, quantization, batching, memory, and serving**. - He says Kimi K3 currently uses roughly **2x the tokens** of GPT-5.6 for the same task, but that gap should narrow as the ecosystem adapts. - His broader point: U.S.-based inference companies will spend heavily optimizing Chinese open-model weights once they are available, similar to how Fireworks, Modal, and Baseten have built large businesses around inference infrastructure.

Related event: Kimi K3 Reshapes Global AI Pricing, Sparking China-US Compute Cost Debate(9 posts)→

Original post →

More from Venture

Venture channel →