OpenAI's Math Model Scores Just 9.3% on 'All the Math We Can Think Of' Bench (372/4000)

basedjensen · x · 2026-10-07

Amid buzz over OpenAI's model-produced math results, a commenter flagged its score on the 'All the Math We Can Think Of Bench': 372 hits out of 4,000 problems (9.3%), with only 3 hours per problem. The poster wants failure analysis and which math problem families are most resistant — a notable contrast between headline theorems and low broad coverage.

Original post →

More from Models

Models channel →