OpenAI reportedly folded open-problem lists into internal math benchmarks, sparking clarification
tszzl · x · 2026-10-07
tszzl and Fields medalist Tim Gittens (@littmath) discussed the provenance of OpenAI's math results. Gittens noted his own small problem list appears to have been added to an internal OpenAI benchmark; tszzl admitted he can't verify how each list got there, guessing researchers wanted the answers rather than a centralized 'solve all open problems' effort, and said he might delete the original post.
Related event: OpenAI Set to Release Hundreds of AI Math Proofs, Angering Mathematicians(20 posts)→
More from Models
- Dev: Claude wins benchmarks but Codex better at doing what I want — Kuprel · 2026-10-07
- DataCamp founder slams OpenAI Codex's 'insane' automatic resets — hugobowne · 2026-10-07
- Developer says 6.1 is underrated: handling GIS automation to 3D reconstruction well — haider1 · 2026-10-07
- HuatuoGPT-3: 27B open medical LLM hits 70.1 on HealthBench, beating GPT-6 Astra — CUHKSZ · 2026-10-07
- Dev Complains Overnight Long-Running Tasks Keep Hitting Usage Limits — willdepue · 2026-10-07
- Gary Marcus on whether frontier LLMs can solve open math problems without symbolic harnesses — GaryMarcus · 2026-10-07