Epoch AI unveils FrontierMath Erdős: pre-release Astra solves only 2 of 68
Chris-MelodyFirst · reddit · 2026-09-04
Epoch AI announces FrontierMath Erdős, a new 68-problem benchmark. A pre-release Astra solved 2 of 68, while GPT-5.6 Sol, GPT-5.5, Claude Fable 5.1 and Claude Fable 5 all failed on every attempt, making it by far the hardest math benchmark for frontier models. Details in Epoch AI's announcement.
More from Research
- Claude's Fermat Proof Passes Independent Rust Verifier: 1,052,234 Declarations, Zero Errors — imjustnewatai · 2026-09-05
- AI-enhanced adaptive virtual screening over billions of molecules lands in Nature Biotech — anshulkundaje · 2026-09-05
- F. Chollet: all AI will converge to symbolic learning as the optimally efficient form — burny_tech · 2026-09-05
- Anthropic model formalizes Fermat's Last Theorem in Lean, 13.4M lines closing 100-theorem benchmark — littmath · 2026-09-05
- GPT-6 Astra hits record 169 on Epoch AI's ECI, sweeping math and continual-learning benchmarks — rohanpaul_ai · 2026-09-05
- SemiAnalysis: Zhipu experiments with Loop Transformer as RL environment building becomes the new bottleneck — ricklamers · 2026-09-05