OpenAI's Math Model Scores Just 9.3% on 'All the Math We Can Think Of' Bench (372/4000)
basedjensen · x · 2026-10-07
Amid buzz over OpenAI's model-produced math results, a commenter flagged its score on the 'All the Math We Can Think Of Bench': 372 hits out of 4,000 problems (9.3%), with only 3 hours per problem. The poster wants failure analysis and which math problem families are most resistant — a notable contrast between headline theorems and low broad coverage.
More from Models
- Is Bel's math edge scale or synthetic data? TeortaxesTex bets on data — teortaxesTex · 2026-10-07
- Fable claims its new release equals roughly five OpenAI Navier-Stokes-level results — willdepue · 2026-10-07
- Reddit user ships 'surgical abliterated' 27B red-team model with zero refusals — Least_Dog_8556 · 2026-10-07
- OpenAI Researcher Surprised AI Lab Math Results So Far All Hold Up — willdepue · 2026-10-07
- OpenAI dots losing to Meta's Muse surprises AI community — BLUECOW009 · 2026-10-07
- Inception launches Mercury Decide on OpenRouter: free structured-decision model doing 14 decisions/sec — StefanoErmon · 2026-10-07