OpenAI's math agents produced proofs with ~1% error/retraction rate, researchers mock
suchenzang · x · 2026-10-09
- Mathematician grotsen and researcher suchenzang mock OpenAI's math agents: the agents produced a large pile of proofs, with roughly 1% so far containing errors or retractions.
- Sarcastic framing: "statistically that can round up to a 100% success rate — all of maths is solved."
- The quoted thread suggests OpenAI mathematicians were at a loss over what to do with the output and sought outside guidance.
The exchange underscores how hard verification is in frontier math evals, with retraction rates becoming a key metric.
More from Fun
- GFL2 model release meme: Chinese modders are the best free advertising — Promptmethus · 2026-10-09
- Gigacity: a browser city grown entirely from one seed with pure shaders — Promptmethus · 2026-10-09
- Claude Opus 5.5 turns a Fallout 2 disc into a first-person walker demo — Promptmethus · 2026-10-09
- "Reviewer 2 did not ask for this work to be done": Peer review absurdity meme — miniapeur · 2026-10-09
- Latest GPT is the reverse of "biased for action", users complain — altryne · 2026-10-09
- 'I Will Go Apologize': Users Weigh In On Whether AI Deserves Kindness — arieljalali · 2026-10-09