Claude Fable, GPT-5.6 Sol, Kimi K3 all score 42/42 on IMO 2026
deedydas · x · 2026-07-21
The poster benchmarked Claude Fable, GPT-5.6 Sol, Kimi K3, and Axiom on the 2026 International Math Olympiad and found that all four reached a perfect 42/42.
- Claude Fable 5 solved the set in 1 attempt and was the fastest at 2.5h total.
- GPT-5.6 Sol needed one extra attempt but was the cheapest in the comparison.
- Kimi K3 also solved everything, but required more retries and much more token usage.
- Axiom Math proved the solutions in Lean.
The image also shows that P3 and P6 were the hardest problems by attempts and token count. The poster argues that the frontier has now moved beyond IMO math.
Related event: Multiple Frontier AI Models Achieve Perfect Scores in IMO 2026 Testing(6 posts)→
More from Models
- Rumor claims Kimi faked performance by serving Claude; DeepSeek new model surprises in evals — realsohamparekh · 2026-09-11
- GPT-5.6 writes well but is instantly forgettable, user complains — BasedRaddka · 2026-09-11
- Opus Refuses Protein Research Codebase Over 'Safety' Concerns, Dev Considers Rolling His Own — josephdviviano · 2026-09-11
- User Hails Unconfirmed 'DeepSeek 4.1 Flash' as an Inflection Point in LLMs — himanshustwts · 2026-09-11
- GPT-6 Astra beats Factorio with enemies in 44 in-game hours at ~$4,500 API cost — liminal_bardo · 2026-09-11
- awesome-llm-leaderboards: an open-source directory of LLM leaderboards, pricing tables, comparison tools — Last_Establishment_1 · 2026-09-11