Epoch AI math benchmark goes from 5% to 98% in under 14 months, now saturated
Jsevillamol · x · 2026-09-11
Epoch AI says its Tier 4 math benchmark is now saturated: the top score was just 5% when it launched on July 11, 2025, and less than 14 months later it sits at 98%. The thread underscores how the exponential in model math capability has yet to hit a ceiling.
Related event: Epoch AI's Tier 4 Math Benchmark Saturated in 14 Months(3 posts)→
More from Models
- Novel Reasoning Effort Control Scheme Analyzed: Graded GRPO Training — stochasticchasm · 2026-09-11
- Forcing models to always max effort is like humans evolving on Adderall, researcher argues — voooooogel · 2026-09-11
- Pushing models to always show 'maximum effort' drags along its corollaries, dev argues — voooooogel · 2026-09-11
- Claude is the distillation target of choice because agentic RL seed data is scarce — teortaxesTex · 2026-09-11
- Why Chinese labs distill from Anthropic: Claude's agent data is the scarce training signal — teortaxesTex · 2026-09-11
- ApprenticeBench: closed model scores 72% vs open Kimi K3 at 18% on real jobs — ysu_nlp · 2026-09-11