Epoch AI Launches FrontierMath Erdős: GPT-6 Astra Scores 3%, All Other Models Zero

haider1 · x · 2026-09-05

Epoch AI announced FrontierMath Erdős, a benchmark of 68 significant unsolved Erdős problems curated by Thomas Bloom and formalized in Lean, where AI systems must prove or disprove each within a fixed compute budget. It addresses curation, verification, and replicability issues informal Erdős-benchmark usage has had (erdosproblems.com lists 1,217 problems, 652 unsolved). GPT-6 Astra scored 3% — the only nonzero result: Fable 5.1, Fable 5, Sol 5.6, and 5.5 all scored 0%. Across additional attempts Astra eventually solved 5 of 68 open problems, but at over $220,000 in compute cost.

Related event: Epoch AI Launches FrontierMath Erdős Benchmark; GPT-6 Astra Is Only Model to Score Above Zero(10 posts)→

Original post →

More from Models

Models channel →