K3 Model Hits Only 26 tps on OpenRouter
nrehiew_ · x · 2026-07-17
According to OpenRouter data, the K3 model currently runs at an inference speed of about 26 tps (tokens per second), only half the speed of Claude Opus.
The poster hopes that upon the official release, inference acceleration frameworks like vllm or sglang will optimize the model to boost operational efficiency.
Related event: Kimi K3 Inference Speed Lags Behind Claude Opus(3 posts)→
More from Models
- Same Echo Maze prompt, three frontier models: all passed visually but shipped the same hidden bug — eyishazyer · 2026-09-11
- Benchmark scores drop from 89% to 19% on new evals — how benchmaxxing breaks leaderboard trust — airesearch12 · 2026-09-11
- ChatGPT tells user their question is too hard and to 'accept dumber answers' — phido3000 · 2026-09-11
- Claude is no longer available for minors as Anthropic rolls out age assurance — Muhammad523 · 2026-09-11
- Developer Building a Unified Leaderboard of All Model Benchmark Scores — airesearch12 · 2026-09-11
- Rumor claims Kimi faked performance by serving Claude; DeepSeek new model surprises in evals — realsohamparekh · 2026-09-11