GPT-6 Astra tops ErdosBench of 226 open math problems, with only 5-10% gain over GPT-5.6
scaling01 · x · 2026-09-08
przchojecki reports GPT-6 Astra xhigh ranked 1st on ErdosBench, a benchmark of 226 open math problems inspired by Erdős problems.
- The leap over GPT-5.6 Sol xhigh is modest: a few more problems settled, stronger scientific writing, less overclaiming
- Measured gain across math-research skills: roughly 5-10%
- The benchmark remains far from saturated; author awaits GPT-7
More from Models
- GPT-6 Astra scores 14% on no-Python MazeBench, 7x Claude Fable 5.1's 2% — rohanpaul_ai · 2026-09-08
- Nat Lambert: open models deserve attention more than frontier lab drama — deliprao · 2026-09-08
- Hugging Face collection confirms Nex-N2.5 open models up to 1.6T params are live — AdinaYakup · 2026-09-08
- Nex-AGI drops three Apache 2.0 agentic models, topping out at 1.6T params — AdinaYakup · 2026-09-08
- DeepSeek post-training head Wu Yu says V4.1 will come with a price cut — teortaxesTex · 2026-09-08
- DeepSeek rolls back idle-period input pricing, output still 2x pre-hike levels — teortaxesTex · 2026-09-08