Team enlists IOI/IMO/ICPC finalists to test whether AI judges can grade math proofs, using 534 borderline submissions

karinanguyen · x · 2026-10-03

After its first math contest on Repovive, the team worked with IOI, IMO and ICPC finalists to evaluate their AI judge's grading ability.

The core question: can models reliably understand and check challenging mathematical proofs? The author notes accuracy is higher across the full set of contest submissions.

Original post →

More from Models

Models channel →