Kimi K3 ranks second on AA-Briefcase but costs $10.57 and takes 56 minutes per task
ArtificialAnlys · x · 2026-07-22
Artificial Analysis says Kimi K3 ranks second on its AA-Briefcase agentic knowledge-work benchmark, but is expensive and slow to run.
- Kimi K3 scores 1543 Elo on AA-Briefcase, behind only Claude Fable 5 (1574) and far above Kimi K2.6 (816).
- The model reaches 57 on the Artificial Analysis Intelligence Index, roughly comparable to models like Opus 4.8 and GPT-5.5.
- Despite the strong score, it averages $10.57 per task, about 10× Kimi K2.6, with 83 turns per task and an average runtime of 56.4 minutes.
- The report says K3 is strong on objective and analytical work, but weaker on presentation quality than some rivals.
Related event: Kimi K3 Scores High on AA-Briefcase but Incurs High Costs and Long Runtimes(5 posts)→
More from Models
- Elon Musk pushes users to try Grok’s speech-to-text and Build mode — elonmusk · 2026-07-22
- Which labs can mount a model comeback? DeepMind slipping out of the top 10 would be the joke — teortaxesTex · 2026-07-22
- Vision model reads symbol text and answers without tools — john__allard · 2026-07-22
- Gemini Found More Sycophantic Than Doubao in Recent Tests — oran_ge · 2026-07-22
- Trick Discovered: Typing Arabic in English Bypasses Limits in o3 — willdepue · 2026-07-22
- Users are switching GPT-5.6 variants to dodge cybersecurity request blocks — ivan_bezdomny · 2026-07-22