GPT-5.6 Sol Tops ErdosBench

i_dg23 · x · 2026-07-19

GPT-5.6 Sol took first place on **ErdosBench**, solving 78 out of 226 research-level problems, significantly higher than the 55 solved by GPT-5.5 xhigh. The chart also provides detailed statistics across models, including coverage, number of solved problems, strong conclusions/counterexamples/citations, and post-review notes: GPT-5.6 Sol is "the best fully audited run so far"; GPT-5.5 xhigh is generally the best but still has single-model controversies; Kimi K2.7 Code is creative but missing some lines; GLM-5.2 is strong but unbalanced; and Claude Opus 4.8 max performs better in reviewing and partial depth.

Original post →

More from Models

Models channel →