Kimi K3 ranks second on AA-Briefcase but costs $10.57 and takes 56 minutes per task
ArtificialAnlys · x · 2026-07-22
Artificial Analysis says Kimi K3 ranks second on its AA-Briefcase agentic knowledge-work benchmark, but is expensive and slow to run.
- Kimi K3 scores 1543 Elo on AA-Briefcase, behind only Claude Fable 5 (1574) and far above Kimi K2.6 (816).
- The model reaches 57 on the Artificial Analysis Intelligence Index, roughly comparable to models like Opus 4.8 and GPT-5.5.
- Despite the strong score, it averages $10.57 per task, about 10× Kimi K2.6, with 83 turns per task and an average runtime of 56.4 minutes.
- The report says K3 is strong on objective and analytical work, but weaker on presentation quality than some rivals.
Related event: Kimi K3 Ranks Second on AA-Briefcase but with High Costs and Long Runtimes(6 posts)→
More from Models
- Bug Hunt Bench ranks frontier coding models on 105 planted real-repo bugs — PawelHuryn · 2026-09-11
- GPT-6 Astra beats Factorio with enemies in 44 in-game hours at ~$4,500 API cost — liminal_bardo · 2026-09-11
- 105 hidden bugs, 2 repos: DeepSeek V4.1 Flash fixes 24 at $1.80 vs Opus 5's 27 at $51.33 — ChartsJournalX · 2026-09-11
- awesome-llm-leaderboards: an open-source directory of LLM leaderboards, pricing tables, comparison tools — Last_Establishment_1 · 2026-09-11
- Anthropic claims it works to keep eval environments unidentifiable to models — MaxKannen · 2026-09-11
- Nex N2.5 Pro released on Hugging Face with 407GB of weights — jinnyjuice · 2026-09-11