Kimi K3 Ties for 5th in Coding, Beats GPT-5.6 in Value
scaling01 · x · 2026-07-18
In the Artificial Analysis Coding Agent Index, Kimi K3 scored 57, tying for 5th place with models like GPT-5.6 Terra and leading Opus 4.8. In specific evaluations, it delivered outstanding results in Terminal-Bench, DeepSWE, and others. Furthermore, the average cost per task for K3 is only $3.18—55% cheaper than GPT-5.6 Sol max—demonstrating exceptional cost-effectiveness for frontier coding.
More from Models
- Developer Building a Unified Leaderboard of All Model Benchmark Scores — airesearch12 · 2026-09-11
- Rumor claims Kimi faked performance by serving Claude; DeepSeek new model surprises in evals — realsohamparekh · 2026-09-11
- GPT-5.6 writes well but is instantly forgettable, user complains — BasedRaddka · 2026-09-11
- Opus Refuses Protein Research Codebase Over 'Safety' Concerns, Dev Considers Rolling His Own — josephdviviano · 2026-09-11
- User Hails Unconfirmed 'DeepSeek 4.1 Flash' as an Inflection Point in LLMs — himanshustwts · 2026-09-11
- Terminal Bench v4: GLM-5.3 Leads at 41.9%, Kimi-K3 Underwhelms at 12.6% — Ok_Warning2146 · 2026-09-11