Claude Opus 5 reportedly scores 42/42 on IMO 2026 without tools
surmenok · x · 2026-07-26
Claude Opus 5 is reported to have solved all 42 IMO 2026 problems without an agent harness or tools, reaching a gold-medal-level 42/42. The poster argues this means the benchmark is now fully saturated and notes that even with four independent solutions, all answers were correct.
The same thread adds that these models are now roughly five standard deviations above the average human, framing the result as a sign of how far frontier models have pushed olympiad-style math evaluation.
Related event: Rumor Claims Claude Opus 5 Achieves Perfect 2026 IMO Score Without Tools(5 posts)→
More from Models
- Looking for a classifier of software engineering task shapes to pick models per task — StewartalsopIII · 2026-09-11
- DeepSeek V4 Pro API to continue after Sept 2026, billing unchanged — teortaxesTex · 2026-09-11
- DeepSeek V4.1 Flash Hits 98% of GPT-6 Astra's Score at 1.4% of the Cost in Third-Party Benchmark — ayushtweetshere · 2026-09-11
- TheZvi Polls: Has Your Coding Model Choice Changed Since Fable 5.1 and Astra? — TheZvi · 2026-09-11
- antirez Weighs In on Anthropic Banning Minors From Using Claude — antirez · 2026-09-11
- Engram's random reads don't suit SSDs; CPU-memory over NVLink could serve all 72 GPUs — bookwormengr · 2026-09-11