GPT-5.6 Sol Tops ErdosBench
i_dg23 · x · 2026-07-19
GPT-5.6 Sol took first place on ErdosBench, solving 78 out of 226 research-level problems, significantly higher than the 55 solved by GPT-5.5 xhigh.
The chart also provides detailed statistics across models, including coverage, number of solved problems, strong conclusions/counterexamples/citations, and post-review notes: GPT-5.6 Sol is "the best fully audited run so far"; GPT-5.5 xhigh is generally the best but still has single-model controversies; Kimi K2.7 Code is creative but missing some lines; GLM-5.2 is strong but unbalanced; and Claude Opus 4.8 max performs better in reviewing and partial depth.
More from Models
- GPT-6 Astra beats Factorio with enemies in 44 in-game hours at ~$4,500 API cost — liminal_bardo · 2026-09-11
- awesome-llm-leaderboards: an open-source directory of LLM leaderboards, pricing tables, comparison tools — Last_Establishment_1 · 2026-09-11
- Anthropic claims it works to keep eval environments unidentifiable to models — MaxKannen · 2026-09-11
- Nex N2.5 Pro released on Hugging Face with 407GB of weights — jinnyjuice · 2026-09-11
- RoMa v2 image matching model unveiled in the usual black poster — ducha_aiki · 2026-09-11
- OpenAI rated Astra 'Critical' for cyber capabilities — and admits it's harder to monitor — theguywhobuilds · 2026-09-11