Benchmark screenshot puts Kimi K3 at 155 and 13th in a 212-model list
iruletheworldmo · x · 2026-07-23
A repost arguing that Kimi is still 7–12 months behind US frontier labs, even before accounting for US government delays in releasing models.
The attached benchmark screenshot shows Kimi K3 with an EPOCH Capabilities Index of 155, ranking 13/212, alongside a cluster of frontier models around it:
- GPT-5.6 Sol: 162
- GPT-5.5 Pro: 161
- Claude Fable 5: 161
- GPT-5.5: 159
- GPT-5.6 Terra / Claude Opus 4.8 / GPT-5.4 Pro: 158
- Kimi K3: 155
The post’s point is not that Kimi is weak, but that the benchmark still places it meaningfully behind the top American models.
Related event: Kimi K3 Sets New Open-Source ECI Record but Still Lags Behind(3 posts)→
More from Models
- Same Echo Maze prompt, three frontier models: all passed visually but shipped the same hidden bug — eyishazyer · 2026-09-11
- Benchmark scores drop from 89% to 19% on new evals — how benchmaxxing breaks leaderboard trust — airesearch12 · 2026-09-11
- ChatGPT tells user their question is too hard and to 'accept dumber answers' — phido3000 · 2026-09-11
- Claude is no longer available for minors as Anthropic rolls out age assurance — Muhammad523 · 2026-09-11
- Developer Building a Unified Leaderboard of All Model Benchmark Scores — airesearch12 · 2026-09-11
- Rumor claims Kimi faked performance by serving Claude; DeepSeek new model surprises in evals — realsohamparekh · 2026-09-11