Multiple Frontier AI Models Achieve Perfect Scores in IMO 2026 Testing

Recently, several frontier AI models achieved historic breakthroughs in the International Mathematical Olympiad (IMO). According to data released by evaluator @deedydas, four models—Claude Fable 5, GPT-5.6 Sol, Kimi K3, and Axiom—all achieved a perfect score of 42/42 on the IMO 2026 test. This performance marks a significant leap in AI's ability to solve highly complex mathematical problems.

Key Details and Cost Comparison

@deedydas pointed out that this is the first time the IMO has been "completely solved." For comparison, last year's best performers—an unreleased Gemini Deep Think and an experimental OpenAI model—only scored 35/42. In this year's tests, three public models successfully solved the problems for a cost of just $10 to $50. Furthermore, Claude Fable 5 solved the problems using only a single attempt. The testing process did not use public web search; all run links, repository addresses, and Lean proofs have been open-sourced for reproducibility and verification.

Reactions and Skepticism

While these groundbreaking scores attracted community attention, they also brought scrutiny regarding the rigor of the tests. @deedydas emphasized that all process data is available in the code repository, though actual results might vary depending on the number of attempts due to differences in testing conditions. @bdsqlsz, who reshared the data, used it as a horizontal reference point to compare the gaps in model performance, cost, and overall capabilities.

2026-07-21 ~ 2026-07-21 · 5 related posts