Kimi K3 ranks second on AA-Briefcase, but each task costs about $10.57
airesearch12 · x · 2026-07-22
Kimi K3 is being highlighted as a strong agentic-work model: it ranks second on AA-Briefcase with an Elo of 1543, just behind Claude Fable 5 at 1574.
The catch is cost. The benchmark estimate puts Kimi K3 at about $10.57 per task, roughly 10× Kimi K2.6 and even above Opus 4.8 on this workload.
The post also notes:
- Kimi K2.6 jumped from 816 to 1543 for K3, a +727 Elo gain
- K3 reached 51% rubric pass rate, versus 56% for Fable 5
- Its analytical Elo was 1754, nearly matching Fable 5’s 1744
Related event: Kimi K3 Ranks 2nd on AA-Briefcase but with High Cost and Latency(6 posts)→
More from Models
- Bloomberg: Sam Altman to brief U.S. officials on OpenAI’s next wave of AI models — pstAsiatech · 2026-07-22
- Claude saves a four-point memory rule after a paper-reading math error — bookwormengr · 2026-07-22
- Teknium jokes that GPT-5.6 SOL “must be AGI” because it never stops pursuing the task — Teknium · 2026-07-22
- Solar Open 2 is a 250B open-weight model aimed at agentic workloads — jacek2023 · 2026-07-22
- Upstage’s Solar Open 2 posts strong coding and reasoning scores against DeepSeek-V4-Flash — Lucidstyle · 2026-07-22
- AutoLab benchmark shows frontier models win long-horizon tasks by persisting, not guessing — rohanpaul_ai · 2026-07-22