FULL STORY
FrontierMath Erdős Debuts as GPT-6 Astra Sets New Records
Epoch AI launched the FrontierMath Erdős math benchmark, with GPT-6 Astra becoming the first model to score above zero. GPT-6 Astra then set a new Epoch Capability Index record at 169 points.
2026-09-04 ~ 2026-09-07 · 2 episodes · 13 posts
Episode 1 · Epoch AI Launches FrontierMath Erdős; GPT-6 Astra First to Break Zero (2026-09-04, 10 posts)
Epoch AI, together with mathematician and erdosproblems.com founder Thomas Bloom (University of Manchester), released a new math benchmark, FrontierMath Erdős: 68 problems selected from Erdős problems still unsolved as of August 2026, requiring AI systems to produce Lean-formalized, automatically verifiable proofs within a fixed compute budget. GPT-6 Astra became the first public model with a solve rate above 0%, officially solving 2/68 (about 3%); all other top models scored zero.
Confirmed
- The benchmark was built by Epoch AI with Thomas Bloom: 68 unsolved Erdős-type problems answered via Lean-formalized proofs within a fixed compute budget.
- Among five top models tested, GPT-6 Astra was the only nonzero scorer: 2 of 68 problems (about 3%); GPT-5.6 Sol, GPT-5.5, Claude Fable 5.1 and others all scored 0.
- The 3% figure circulating on Reddit and via DrSingularity is consistent with the official 2/68 count.
- A Reddit user posted a chart claiming GPT-6 Astra leads SOL, Fable and Spark on a coding-agent index and ranks first on math benchmarks, but the specific scores in the image could not be verified.
Unconfirmed
- Chinese outlet QbitAI reported Astra "uniquely solved 5" Erdős problems, conflicting with the official 2/68 figure; the official release should be treated as authoritative.
Why it matters
- As @Jsevillamol relayed from Greg Burnham's idea, testing AI on Erdős-level open problems with Lean auto-verification directly probes whether models can produce formal proofs humans have not yet completed, giving a quantifiable yardstick for genuine AI mathematical discovery. Even 2/68 marks the first time a public model breaks zero on such a benchmark.
- Astra becomes first public model to solve any curated hard Erdős problems with Lean proofs — Jsevillamol · 2026-09-04
- Epoch AI launches FrontierMath Erdős benchmark; Astra solves 2 of 68 unsolved problems — littmath · 2026-09-04
- Epoch AI launches FrontierMath Erdős benchmark; GPT-6 Astra first to crack two problems — Jsevillamol · 2026-09-04
- Epoch AI launches FrontierMath Erdős benchmark of 68 unsolved problems; Astra tops it at 2/68 — keviv9 · 2026-09-04
- GPT-6 Astra Only Non-Zero Scorer on 68 Unsolved Erdős Problems, Solves 5 — 量子位 · 2026-09-04
- Astra reportedly leads coding agent index over SOL, Fable, and Spark — XCxBigDong69XCx · 2026-09-04
- Epoch AI unveils FrontierMath Erdős: pre-release Astra solves only 2 of 68 — Chris-MelodyFirst · 2026-09-04
- GPT-6 Astra scores 3% on FrontierMath Erdős benchmark while every other tested model scores 0% — Every_Foundation5197 · 2026-09-05
- GPT-6 Astra reportedly scores 3% on FrontierMath Erdős, spending $220K in compute — haider1 · 2026-09-05
- Epoch AI Launches FrontierMath Erdős: GPT-6 Astra Scores 3%, All Other Models Zero — haider1 · 2026-09-05
Episode 2 · GPT-6 Astra Sets Epoch AI ECI Record at 169 (2026-09-05, 3 posts)
GPT-6 Astra scores 169 on Epoch AI's ECI, beating the previous record of 163 and setting new marks across math, continual learning and puzzle benchmarks, with FrontierMath Tier 4 saturating at 98%.
- GPT-6 Astra hits record 169 on Epoch AI's ECI, sweeping math and continual-learning benchmarks — rohanpaul_ai · 2026-09-05
- GPT-6 Astra lifts Epoch capability index 6 points, saturates FrontierMath Tier 4 at 98% — eyishazyer · 2026-09-06
- GPT-6 Astra Sets New ECI Record at 169, Beating Prior Best of 163 — burny_tech · 2026-09-07