OpenAI criticized for claiming 100 open math problems solved without disclosing the total attempted
burny_tech · x · 2026-09-22
marcusgnt publicly called out OpenAI's eval reporting: the company claims its model solved 100 open math problems, but never disclosed how many were attempted. He argues reporting the success rate is basic scientific practice and essential context for judging model ability, urging OpenAI to publish that simple number for the math community's sake. The main post is a repost endorsing this critique.
More from Models
- Grok 4.7 launches on Cursor and API, topping coding benchmarks at half the price — FinanceYF5 · 2026-09-22
- Grok 4.7 launches at same pricing: Terminal-Bench doubles to 38%, 500K context kept — FinanceYF5 · 2026-09-22
- Claude Status: Elevated Errors Reported for Multiple Models — corvad · 2026-09-22
- tenobrus and antirez pour cold water on Jev: demos are inflated and far from functional — burny_tech · 2026-09-22
- Musk confirms Grok went from outside top 10 to top 3 in 90 days, Grok 4.8 next — elonmusk · 2026-09-22
- Anthropic investigates elevated errors across Claude Mythos 5.1, Fable 5.1 and Opus 5 — ClaudeAI-mod-bot · 2026-09-22