Multiple Frontier AI Models Achieve Perfect Scores in IMO 2026 Testing

Recently, several frontier AI models achieved historic breakthroughs in International Mathematical Olympiad (IMO) testing. According to independent test data released by former Google engineer @deedydas, Claude Fable 5, GPT-5.6 Sol, and Kimi K3 all achieved perfect scores of 42/42 in a horizontal evaluation of all 6 problems from IMO 2026. This performance marks a significant leap in AI's ability to solve highly difficult mathematical problems.

Key Details and Cost Comparison

@deedydas pointed out that this is the first time the IMO has been "completely solved." In comparison, last year's best performers were an unreleased Gemini Deep Think and an OpenAI experimental model, which only scored 35/42. In this year's test, 3 public models successfully solved the problems for a cost of only $10 to $50 (according to a report by @新智元, the lowest cost was merely $20). Additionally, Claude Fable 5 used only 1 attempt during the problem-solving process. The testing did not use public web search, and all run links, repository addresses, and Lean proofs have been open-sourced for verification.

Reactions and Doubts

These breakthrough results have attracted community attention, but they have also been accompanied by scrutiny regarding the rigor of the tests. @deedydas emphasized that all process data is placed in the code repository, but due to differences in testing conditions, the actual results might be affected by the number of attempts. @bdsqlsz, who reposted the data, also used it as a horizontal reference for comparing the gaps in model performance, cost, and comprehensive capabilities.

2026-07-21 ~ 2026-07-22 · 6 related posts

Primary sources