Team enlists IOI/IMO/ICPC finalists to test whether AI judges can grade math proofs, using 534 borderline submissions
karinanguyen · x · 2026-10-03
After its first math contest on Repovive, the team worked with IOI, IMO and ICPC finalists to evaluate their AI judge's grading ability.
- They selected 534 suspicious or borderline submissions for review
- Established a shared grading standard and created expert reference judgments
- Compared models under different grading instructions against those expert judgments
The core question: can models reliably understand and check challenging mathematical proofs? The author notes accuracy is higher across the full set of contest submissions.
More from Models
- Steve Yegge: Two Weeks With Opus 5.5 — Precision Rivals Fable, Recall Trails on Open-Ended Tasks — Steve_Yegge · 2026-10-03
- OpenAI's Codex global reset appears to skip Business accounts, support suggests buying credits — AdventurousFeeling19 · 2026-10-03
- Mystery 'iguana_necktie' Field Spotted in Anthropic Usage API — bytebot · 2026-10-03
- Insider teases 'new SSI model,' calling Ilya 'a truly remarkable human being' — iruletheworldmo · 2026-10-03
- Exploit Bench results are highly harness-sensitive: GLM 5.3 Flash beats 4.1, best in Claude Code — teortaxesTex · 2026-10-03
- Reddit user says Opus 5.5 burns weekly cap at 300M tokens, down from 1-2B before — Chemical-Ad-7982 · 2026-10-03