FrontierMath Tier 4, once the hardest math benchmark, falls after 1.5 years
Jsevillamol · x · 2026-09-11
@DotCSV reports that FrontierMath Tier 4 — until recently the most demanding math benchmark — has been solved after roughly a year and a half. It's another landmark showing how fast frontier models' mathematical reasoning has climbed: once again, the hardest ceiling has been broken and the benchmark retires.
Related event: GPT-6 Astra Solves Final Problem as FrontierMath Tier 4 Fully Saturated(6 posts)→
More from Models
- Live-reading Anthropic's 1022-page model transcript: 80 pages obsessing over hCaptcha frogs — voooooogel · 2026-09-11
- Gradio founder asks: why would anyone still pay API prices for closed models? — Sentdex · 2026-09-11
- Unverified DeepSeek-V4.1-Flash report: 552B MoE slashing KV cache for million-token agent workloads — burkov · 2026-09-11
- DeepSeek ships research artifacts, not products — explaining its eval gaps — teortaxesTex · 2026-09-11
- Claude responds to 'how do you know humans are real?' with simulation musings — vishalmisra · 2026-09-11
- Frontier models like GPT-6 Astra excel at one thing: relentlessly pursuing verifiable objectives — daniel_mac8 · 2026-09-11