Kimi K3 Scores High on AA-Briefcase but Incurs High Costs and Long Runtimes
Artificial Analysis released a comprehensive evaluation of Kimi K3's performance on the AA-Briefcase agentic knowledge work benchmark. While Kimi K3 achieved frontier-level overall scores, it did so at the expense of operational efficiency and cost, exhibiting a significant trade-off between strengths and weaknesses.
Core Scores and Capability Imbalance
Kimi K3 secured an overall score of 1543 Elo in the AA-Briefcase test, ranking second and approaching the top-tier models on the leaderboard. In specific dimensions, its analysis quality performed strongly at 1754 Elo, roughly matching Claude's level; however, its presentation quality was notably weaker than its own analytical capabilities, showing an unbalanced skill set.
Operational Cost and Efficiency Trade-offs
Despite the impressive score, the operational cost of running Kimi K3 is exceptionally high. The average task takes 56.4 minutes, making it the most time-consuming model on the leaderboard. The average cost per task also reached $10.57 (second only to the Claude Sonnet 5 max version at $14.43 among the displayed models). The primary reason for the surge in cost and time is that the model requires an average of 83 rounds of interaction per task, generating approximately 120,000 output Tokens.
2026-07-22 ~ 2026-07-22 · 5 related posts
- [source] Kimi K3 ranks second on AA-Briefcase but costs $10.57 and takes 56 minutes per task — ArtificialAnlys · 2026-07-22
- [source] Kimi K3 nears the top of AA-Briefcase, but its presentation score trails — ArtificialAnlys · 2026-07-22
- Kimi K3 scores near the top, but costs $10.57 per task on AA-Briefcase — ArtificialAnlys · 2026-07-22
- [source] Kimi K3 posts top AA-Briefcase results, but costs $10.57 per task and 56.4 minutes — ArtificialAnlys · 2026-07-22
- Kimi K3 Evaluation: 56 Mins Per Task, Token Usage Spikes — ArtificialAnlys · 2026-07-22