Kimi K3 Inference Speed Revealed, Optimization Needed

nrehiew_ · x · 2026-07-17

OpenRouter data shows that Kimi K3 currently runs at about 26 tps, only half the speed of Claude 3 Opus. The author speculates that speculative decoding (spec decoding) isn't enabled yet and expects significant speed improvements once optimized using community frameworks like vllm or sglang.

Related event: Kimi K3 Inference Speed Lags Behind Claude Opus(3 posts)→

Original post →

More from Models

Models channel →