K3 Model Hits Only 26 tps on OpenRouter
nrehiew_ · x · 2026-07-17
According to OpenRouter data, the K3 model currently runs at an inference speed of about 26 tps (tokens per second), only half the speed of Claude Opus.
The poster hopes that upon the official release, inference acceleration frameworks like vllm or sglang will optimize the model to boost operational efficiency.
Related event: Kimi K3 Inference Speed Lags Behind Claude Opus(3 posts)→
More from Models
- NVIDIA says Nemotron 3 Ultra scored 30/42 on the 2026 IMO problems — NVIDIAAI · 2026-07-22
- OpenAI is reportedly briefing U.S. lawmakers on its next model family — kimmonismus · 2026-07-22
- Muse Spark 1.1 lands at 1495 on Text Arena with standout agentic-coding price performance — ycombinator · 2026-07-22
- Advanced AI Models Are Becoming Impossible to Plug and Play — emollick · 2026-07-22
- Google Gemini's AI Problem: No Leading Model for Core Workloads — bindureddy · 2026-07-22
- Model Offers 1M Token Context Window at Just $0.33/1M Tokens — MickeySteamboat · 2026-07-22