Epoch AI Launches FrontierMath Erdős: GPT-6 Astra Scores 3%, All Other Models Zero
haider1 · x · 2026-09-05
Epoch AI announced FrontierMath Erdős, a benchmark of 68 significant unsolved Erdős problems curated by Thomas Bloom and formalized in Lean, where AI systems must prove or disprove each within a fixed compute budget. It addresses curation, verification, and replicability issues informal Erdős-benchmark usage has had (erdosproblems.com lists 1,217 problems, 652 unsolved). GPT-6 Astra scored 3% — the only nonzero result: Fable 5.1, Fable 5, Sol 5.6, and 5.5 all scored 0%. Across additional attempts Astra eventually solved 5 of 68 open problems, but at over $220,000 in compute cost.
More from Models
- Yoav Goldberg: capabilities once dependent on the harness are now baked into the model — yoavgo · 2026-09-05
- Yoav Goldberg: Ark's harness was simply bad, and OpenAI's fix was obvious — yoavgo · 2026-09-05
- Anthropic Fable 5.1 vs OpenAI Astra: analyst teases a clear winner — dylan522p · 2026-09-05
- Reviewer: OpenAI's GPT-6-Astra finally 'gets what you mean,' with Fable-level intelligence and real gains in game dev — pvncher · 2026-09-05
- Meta ships Muse Spark 1.3 with max reasoning, pitching frontier performance at non-frontier prices — AIatMeta · 2026-09-05
- LLMs are now making up words that don't exist, not just jargon — StewartalsopIII · 2026-09-05