Kimi K3 throughput jumps from 19 to 49 tok/s on OpenRouter
cedric_chee · x · 2026-07-28
Kimi K3’s throughput on OpenRouter reportedly jumped from 19 tok/s to 49 tok/s.
The poster also notes a behavioral quirk: K3 seems to spend a surprising amount of its reasoning budget second-guessing itself, repeatedly burning tokens on phrases like “Wait, actually...”. The comparison is tied to an OpenRouter performance readout showing K3 versus GLM-5.2.
More from Infra
- Kimi K3 serving stack reaches 423 tok/s after DSpark draft-model tuning — ying11231 · 2026-07-28
- Cloud and AI prices may keep rising under quarter-on-quarter growth pressure — DavidLinthicum · 2026-07-28
- Falling Inference Compute Costs Could Make 'Vibe Hacking' Very Cheap — joshua_saxe · 2026-07-28
- llama.cpp adds Kimi-K3 text model support for local inference — ilintar · 2026-07-28
- Embedding databases are hurting search, says a reply in the thread — davidmanheim · 2026-07-28
- Apple reclaims the top market-cap spot, passing Nvidia at $4.938T — mitchdeg · 2026-07-28