Kimi K3 Ranks Second on AA-Briefcase but with High Costs and Long Runtimes
Artificial Analysis released a multi-dimensional evaluation of Kimi K3's performance on the AA-Briefcase agentic knowledge work benchmark. The assessment shows that while Kimi K3 achieves cutting-edge results in total score, it pays a massive price in operational efficiency and cost, revealing a significant trade-off.
Confirmed
Overall Score and Capability Imbalance: Kimi K3 achieved a total score of 1543 Elo in the AA-Briefcase test, ranking second and closely approaching the top-ranked Claude Fable 5 at 1574 Elo. In specific dimensions, its analysis quality is notably strong, reaching 1754 Elo, which is roughly equivalent to Claude's level; however, its presentation quality is significantly weaker than its own analytical capabilities, showing an imbalanced skill profile.
Operational Cost and Efficiency Penalty: Despite the impressive score, the operational cost of running Kimi K3 is exceptionally high. The average time per task reaches 56.4 minutes, placing it in the longest-running tier on the list. The average cost per task is $10.57 (second only to the $14.43 Claude Sonnet 5 max version among displayed models). The primary cause for this surge in cost and time is that the model requires an average of 83 interaction rounds per task, generating approximately 120,000 output tokens.
Why it matters
This evaluation highlights a typical trade-off faced by current large language models when pursuing high scores in complex agentic tasks: maximizing interaction and reasoning frequency can elevate final task quality, but it inevitably triggers massive token consumption and extended response times. For practical applications, these steep financial and time costs will directly impact the model's commercial viability in real-world business scenarios.
2026-07-22 ~ 2026-07-22 · 6 related posts
- Episode 1: Kimi K3 Tops Frontend Web App Arena with Enhanced English Skills(2026-07-21, 3 posts)
- Episode 2: Kimi K3 Jumps to 4th on Agent Arena Leaderboard(2026-07-21, 6 posts)
- Episode 3: Moonshot Releases 2.8 Trillion Parameter Open-Weight Model Kimi K3(2026-07-21, 8 posts)
- Episode 4: Kimi K3 Matches Fable 5 in SWE Benchmarks at a Third of the Cost(2026-07-21, 7 posts)
- Episode 5: Kimi K3 Sets New Open-Source ECI Record but Still Lags Behind(2026-07-22, 3 posts)
- Episode 6: Kimi K3 Accused of Gaming Benchmarks Instead of Solving Problems(2026-07-22, 2 posts)
- Episode 7: Kimi K3 Ranks Second on AA-Briefcase but with High Costs and Long Runtimes(2026-07-22, 6 posts)
- Episode 8: Kimi K3 Enters Top-Tier AI Model Ranks in Benchmark Tests(2026-07-22, 4 posts)
- Episode 9: Kimi K3 shifts attention from scale to architecture(2026-07-27, 25 posts)
- Episode 10: TokenSpeed Enables Kimi K3 Support on NVIDIA and AMD Platforms(2026-07-27, 2 posts)
- Episode 11: Moonshot's Kimi K3 Launches on Nebius with 1M Context(2026-07-27, 3 posts)
- Episode 12: SGLang Day-0 Support for Kimi K3 Boosts Throughput to 423 tok/s(2026-07-28, 8 posts)
- Episode 13: Moonshot's Kimi K3 Flagship Model Launches on Together AI(2026-07-28, 12 posts)
- Episode 14: Kimi K3 Max Tops Multiple Arena Leaderboards, Open-Source Model Rivals Proprietary(2026-07-28, 11 posts)
- Episode 15: Fireworks Test: Kimi K3 Matches Opus 5 Quality at Fraction of Cost(2026-07-28, 5 posts)
- Episode 16: Kimi K3 Impresses in Early Benchmarks, Sparking Buzz(2026-07-28, 3 posts)
- Episode 17: Local Kimi K3 Beats Cloud Models in 3D Physics Generation Test(2026-07-28, 5 posts)
- Episode 18: Deep Dive into Kimi K3 Tech Report: Engineering Synergy Drives State-of-the-Art Performance(2026-07-28, 21 posts)
- Episode 19: Moonshot AI's Kimi K3 Launches in Japan(2026-07-28, 2 posts)
- Episode 20: Kimi K3 Passes Compound Benchmark Amid Cost Efficiency Concerns(2026-07-29, 2 posts)
Primary sources
- Kimi K3 ranks second on AA-Briefcase but costs $10.57 and takes 56 minutes per task — ArtificialAnlys ·
- Kimi K3 posts top AA-Briefcase results, but costs $10.57 per task and 56.4 minutes — ArtificialAnlys ·
- Kimi K3 nears the top of AA-Briefcase, but its presentation score trails — ArtificialAnlys ·
- [source] Kimi K3 ranks second on AA-Briefcase but costs $10.57 and takes 56 minutes per task — ArtificialAnlys · 2026-07-22
- [source] Kimi K3 nears the top of AA-Briefcase, but its presentation score trails — ArtificialAnlys · 2026-07-22
- Kimi K3 scores near the top, but costs $10.57 per task on AA-Briefcase — ArtificialAnlys · 2026-07-22
- [source] Kimi K3 posts top AA-Briefcase results, but costs $10.57 per task and 56.4 minutes — ArtificialAnlys · 2026-07-22
- Kimi K3 Evaluation: 56 Mins Per Task, Token Usage Spikes — ArtificialAnlys · 2026-07-22
- Kimi K3 ranks second on AA-Briefcase, but each task costs about $10.57 — airesearch12 · 2026-07-22