K3 Model Hits Only 26 tps on OpenRouter

nrehiew_ · x · 2026-07-17

According to OpenRouter data, the K3 model currently runs at an inference speed of about 26 tps (tokens per second), only half the speed of Claude Opus.

The poster hopes that upon the official release, inference acceleration frameworks like vllm or sglang will optimize the model to boost operational efficiency.

Related event: Kimi K3 Inference Speed Lags Behind Claude Opus(3 posts)→

Original post →

More from Models

Models channel →