Kimi-K3 Scores 60.4% on ARC-AGI-2 Benchmark
scaling01 · x · 2026-08-01
The newly released Kimi-K3 model achieves a high score of 60.4% on the challenging ARC-AGI-2 benchmark, demonstrating strong abstract reasoning capabilities.
More from Models
- Claude V4-Flash Reported to Have Vision Defects, Frequently Building Workarounds — teortaxesTex · 2026-08-01
- Jeremy Howard on LLM Anti-Jailbreak: Banning Prefilling Drives Users to Open Source — jeremyphoward · 2026-08-01
- OpenAI Removing GPT-5.4 Models from ChatGPT Starting August 31 — OpenAIDevs · 2026-08-01
- Blind Gemini Tries Building Its Own Eyes via Code Instead of Using Vision Subagents — teortaxesTex · 2026-08-01
- Debate: Can a 300B V4 Pro Model Beat a Newly Released 2.8T Model? — scaling01 · 2026-08-01
- DeepSeek V4 Flash is Basically Free, Luna Max Offers Insane Value After 80% Price Cut — Hesamation · 2026-08-01