Epoch AI's Tier 4 Math Benchmark Saturated: 5% to 98% in 14 Months
Jsevillamol · x · 2026-09-11
- Citing official Epoch AI data: when the Tier 4 math benchmark launched on July 11, 2025, the top score was just 5%; less than 14 months later the top score is 98% and the benchmark is considered saturated.
- The retweeter notes AI's math progress felt sudden: LLM-based models couldn't even add three years ago, now they top hard math evals.
- Another sign that benchmark saturation is outpacing eval design.
Related event: Epoch AI's Tier 4 Math Benchmark Saturated in 14 Months(3 posts)→
More from Models
- OpenAI reportedly has a next-gen model significantly more capable than GPT-6 Astra — sjgadler · 2026-09-11
- Power users want a $400 40x tier as 20x AI coding subscriptions fall short — Snoo-29395 · 2026-09-11
- FrontierMath Tier 4, once the hardest math benchmark, falls after 1.5 years — Jsevillamol · 2026-09-11
- Octen Search debuts third on search index at 16.9s per task, $0.058 cost — ArtificialAnlys · 2026-09-11
- ValsAI: Astra first AI to reach Minecraft Nether Fortress in long-horizon eval — scaling01 · 2026-09-11
- Dev: You Can Tell Who Has a Real RL Pipeline Just From Model Outputs — teortaxesTex · 2026-09-11