Kimi K3 Tops Open-Weight Models on ARC-AGI, Rivaling Frontier Models
mhmazur · x · 2026-08-01
Moonshot's open-weight model Kimi K3 demonstrates exceptional performance relative to cost on the ARC-AGI benchmarks. Verified by ARC Prize, K3 scored 60.4% on ARC-AGI-2 ($1.59/task) and 94.5% on ARC-AGI-1 ($0.77/task).
This makes K3 the highest-scoring open-weight model evaluated on both benchmarks, with performance comparable to closed-weight frontier models released just 2-3 months ago, such as Opus 4.8 (low).
More from Models
- Claude V4-Flash Reported to Have Vision Defects, Frequently Building Workarounds — teortaxesTex · 2026-08-01
- Jeremy Howard on LLM Anti-Jailbreak: Banning Prefilling Drives Users to Open Source — jeremyphoward · 2026-08-01
- OpenAI Removing GPT-5.4 Models from ChatGPT Starting August 31 — OpenAIDevs · 2026-08-01
- Blind Gemini Tries Building Its Own Eyes via Code Instead of Using Vision Subagents — teortaxesTex · 2026-08-01
- Debate: Can a 300B V4 Pro Model Beat a Newly Released 2.8T Model? — scaling01 · 2026-08-01
- DeepSeek V4 Flash is Basically Free, Luna Max Offers Insane Value After 80% Price Cut — Hesamation · 2026-08-01