Kimi K3 Tops Open-Weight Models on ARC-AGI, Rivals Claude Opus
mhmazur · x · 2026-08-01
The Kimi K3 model has achieved outstanding benchmark scores on ARC-AGI, living up to its reputation as an open frontier model.
According to verified data from ARC Prize:
- ARC-AGI-2: Scored 60.4% at $1.59/task.
- ARC-AGI-1: Scored 94.5% at $0.77/task.
It is now the highest-scoring open-weight model evaluated on both benchmarks, with its ARC-AGI-2 performance comparable to Claude Opus (low), which was released two months ago.
More from Models
- DeepSeek v4 Flash Reported to Loop and Forget Context in Coding — kwizzle · 2026-08-01
- Local Deployment on DGX Spark: Exploring Upgrades Beyond Qwen 3.5 122B — Voxandr · 2026-08-01
- 25.5 Trillion Tokens a Day: How SuperPods Power Massive RL Training — zephyr_z9 · 2026-08-01
- Benchmarking Agent Harnesses: Kimi K3 Shines, Claude Code Costs 4x More — omarsar0 · 2026-08-01
- Agentic RL Gives LLMs 'Vitality': Models Become Frantic and Engaged — teortaxesTex · 2026-08-01
- Building apps directly on iPhone using ChatGPT showcased — reach_vb · 2026-08-01