Kimi K3 Ranks Second on AA-Briefcase but with High Costs and Long Runtimes

Artificial Analysis released a multi-dimensional evaluation of Kimi K3's performance on the AA-Briefcase agentic knowledge work benchmark. The assessment shows that while Kimi K3 achieves cutting-edge results in total score, it pays a massive price in operational efficiency and cost, revealing a significant trade-off.

Confirmed

Overall Score and Capability Imbalance: Kimi K3 achieved a total score of 1543 Elo in the AA-Briefcase test, ranking second and closely approaching the top-ranked Claude Fable 5 at 1574 Elo. In specific dimensions, its analysis quality is notably strong, reaching 1754 Elo, which is roughly equivalent to Claude's level; however, its presentation quality is significantly weaker than its own analytical capabilities, showing an imbalanced skill profile.

Operational Cost and Efficiency Penalty: Despite the impressive score, the operational cost of running Kimi K3 is exceptionally high. The average time per task reaches 56.4 minutes, placing it in the longest-running tier on the list. The average cost per task is $10.57 (second only to the $14.43 Claude Sonnet 5 max version among displayed models). The primary cause for this surge in cost and time is that the model requires an average of 83 interaction rounds per task, generating approximately 120,000 output tokens.

Why it matters

This evaluation highlights a typical trade-off faced by current large language models when pursuing high scores in complex agentic tasks: maximizing interaction and reasoning frequency can elevate final task quality, but it inevitably triggers massive token consumption and extended response times. For practical applications, these steep financial and time costs will directly impact the model's commercial viability in real-world business scenarios.

2026-07-22 ~ 2026-07-22 · 6 related posts

Full story(20 episodes)→

Primary sources