GPT-5.6 Sol Tops ErdosBench
i_dg23 · x · 2026-07-19
GPT-5.6 Sol took first place on **ErdosBench**, solving 78 out of 226 research-level problems, significantly higher than the 55 solved by GPT-5.5 xhigh. The chart also provides detailed statistics across models, including coverage, number of solved problems, strong conclusions/counterexamples/citations, and post-review notes: GPT-5.6 Sol is "the best fully audited run so far"; GPT-5.5 xhigh is generally the best but still has single-model controversies; Kimi K2.7 Code is creative but missing some lines; GLM-5.2 is strong but unbalanced; and Claude Opus 4.8 max performs better in reviewing and partial depth.
More from Models
- Korean startup says its model scored 44 on AAII and matches DeepSeek V4 Pro — JungWooHa2 · 2026-07-21
- OpenAI’s GPT-6 is predicted to be far more efficient than Fable — bindureddy · 2026-07-21
- Moonshot spotlights Kimi K3 and its API platform — pstAsiatech · 2026-07-21
- Motif 3 Beta lands on Hugging Face as South Korea’s foundation-model race heats up — Secure_Smoke_4280 · 2026-07-21
- Sakana AI’s Fugu-Cyber update tops real-world security benchmarks — SakanaAILabs · 2026-07-21
- Mythos release drama is being compared to o1, with limited rollout and an open-source clone — nptacek · 2026-07-21