Frontier models ace counterexamples but lag on closed-form math, benchmark designer notes
evijit · x · 2026-09-09
OfirPress recalls saying in his benchmark-design talks that the Millennium Prize Problems would make a benchmark too hard — "it would take 100 years to saturate" — which now looks prescient. evijit responds that saturation depends on the metric: frontier models are already good at finding negative examples and disproving conjectures, but not yet at producing closed-form solutions, so that portion of the benchmark could take much longer to saturate.
Related event: Benchmark expert says his millennium-problem example aged too fast(3 posts)→
More from Models
- Apodex 1.1 launches with open-weight 35B mini model for local deployment — mhdfaran · 2026-09-09
- Apodex open-sources FrontierAgent agent framework and ships 35B open-weight Apodex 1.1 mini — mhdfaran · 2026-09-09
- Accusation: a covert industry packages and sells your work identity to AI labs — kevinafischer · 2026-09-09
- Community Releases Uncensored 27B Fine-tune Using Marchenko-Pastur Noise Distillation — EvilEnginer · 2026-09-09
- ChatGPT Astra Won a Pokémon Battle Autonomously — at $25 Per Match — bucketbot91 · 2026-09-09
- Reddit users slam Claude's opaque usage percentages, demand metered billing — NTXL · 2026-09-09