Math professor grades OpenAI's 722 results: mostly B/C level, one D-level shock
khademinori · x · 2026-10-08
OpenAI announced a broad range of new mathematical results from an internal frontier model, consulted with an independent IAS advisory group. London math professor Abhishek Saha sorted major results into A-D levels; his verdict: some A, mostly B or C, exactly one D — the 'quasi Riemann hypothesis' result. Notably, it all came from one unreleased model at 3 hours of thinking per result, though OpenAI says some proofs aren't fully confirmed.
More from Models
- Researchers Put GPT, Claude and Grok Behind the Wheel of a Real Corolla — Only One Succeeded — nordicinst · 2026-10-08
- ChatGPT on Browser Shows 'Capabilities Reduced' Warning — A New Kind of Rate Limit? — jasondeanlee · 2026-10-08
- ChatGPT Work mode vs Codex: same quota, far more tasks done per 5-hour window — sasik520 · 2026-10-08
- Epoch's InnovationEval: AI agents still far from producing real research innovations — Afinetheorem · 2026-10-08
- 113 decision models in 3 weeks: 70 built on Qwen, sub-cent per call — jonathanmalkin · 2026-10-08
- User reports Haiku 5.5 is a major workflow upgrade in screenshot post — Sorcerer12345 · 2026-10-08