Frontier models ace counterexamples but lag on closed-form math, benchmark designer notes

evijit · x · 2026-09-09

OfirPress recalls saying in his benchmark-design talks that the Millennium Prize Problems would make a benchmark too hard — "it would take 100 years to saturate" — which now looks prescient. evijit responds that saturation depends on the metric: frontier models are already good at finding negative examples and disproving conjectures, but not yet at producing closed-form solutions, so that portion of the benchmark could take much longer to saturate.

Related event: Benchmark expert says his millennium-problem example aged too fast(3 posts)→

Original post →

More from Models

Models channel →