Only 372 hits out of 4,000 problems: analysts probe failure rates in OpenAI's math output
basedjensen · x · 2026-10-07
Commenting on OpenAI's 722 math manuscripts, one analyst notes that 372 successful result families out of 4,000 attempted problems seems like a low hit rate — though only about three hours of thinking compute was spent per problem. He suggests it would be far more informative to see the failure cases and which families of results proved most resistant to the model, rather than only the successes.
More from Models
- Is Bel's math edge scale or synthetic data? TeortaxesTex bets on data — teortaxesTex · 2026-10-07
- Fable claims its new release equals roughly five OpenAI Navier-Stokes-level results — willdepue · 2026-10-07
- Reddit user ships 'surgical abliterated' 27B red-team model with zero refusals — Least_Dog_8556 · 2026-10-07
- OpenAI Researcher Surprised AI Lab Math Results So Far All Hold Up — willdepue · 2026-10-07
- OpenAI dots losing to Meta's Muse surprises AI community — BLUECOW009 · 2026-10-07
- Inception launches Mercury Decide on OpenRouter: free structured-decision model doing 14 decisions/sec — StefanoErmon · 2026-10-07