FrontierMath Tier 4 fully solved 14 months after launch; LEAP panel forecasts lag reality
Afinetheorem · x · 2026-09-11
FrontierMath Tier 4—Epoch AI's benchmark of extremely challenging research-level math problems designed to test the limits of AI reasoning—has now been solved entirely, just 14 months after its introduction.
Afinetheorem compares forecasts: the LEAP panel in summer '25 predicted 55% completion on the easier Tier 1-3 by Dec 2027, while on diffusion tasks LEAP was too optimistic (currently 1% solved vs. a 7.3% forecast). He had personally called 80% and 5%.
The takeaway: research-grade benchmarks are being exhausted faster than expert predictions can keep up.
Related event: GPT-6 Astra solves all FrontierMath Tier 4 problems(2 posts)→
More from Models
- Microsoft Patches Record 974 Vulnerabilities, Mostly Found by AI — Distinct-Question-16 · 2026-09-11
- DeepSeek V4.1 Flash tops Vals open-weight index at $0.30 per test, with the smallest skills gap — teortaxesTex · 2026-09-11
- Do You Really Need Flagship Models? Dev Argues Medium Effort Covers 80% of Coding — iamaliveix · 2026-09-11
- OpenAI appears to be quietly rolling out managed Agents on its platform — testingcatalog · 2026-09-11
- 30B Open Model OpenResearcher Beats GPT-4.1 on BrowseComp-Plus — TheZachMueller · 2026-09-11
- Surge AI evals: Claude Fable 5.1 leads at 68.7, Gemini 3.8 Flash jumps 12 points on frontier math — echen · 2026-09-11