GPT-5.6 achieves ~10% success rate on 3,300 open math problems
littmath · x · 2026-08-25
Sasho revisited a list of open math problems from a decade-old workshop to test AI capabilities, specifically using GPT-5.6 Sol xhigh.
- The model was tested on 3,300 open problems, resulting in roughly 170 counterexamples and 170 proofs.
- This represents a completion rate of about 10%.
- The difficulty of open problems varies significantly, with current AI mostly solving trivialities.
More from Models
- Ox Alpha Model Sparks Speculation: Possible Google Open Weights Bet — zakelfassi · 2026-08-25
- Creative Writing Leaderboard Updates with GLM-5.3 and Gemini-3.7-flash — sam_paech · 2026-08-25
- Open-source AI token usage grows 100x, predicted to overwhelm closed source — bindureddy · 2026-08-25
- Does it exist: A leaderboard for Speech Recognition / STT? — Elibroftw · 2026-08-25
- Open source models rapidly closing the gap with closed source AI — iScienceLuvr · 2026-08-25
- Newer frontier models aren't always better; major labs have all shipped regressions — bindureddy · 2026-08-25