Ofir Press: building good coding benchmarks is about to get much harder

OfirPress · x · 2026-10-07

In a follow-up, Ofir Press argues that building good coding benchmarks was always hard and is about to get harder: it used to be enough to pick progressively harder tasks humans can do, but in the emerging super-human stage, defining new tasks becomes far more challenging.

Related event: OpenAI Math Results Stun Researchers as Superhuman Era Nears(10 posts)→

Original post →

More from AGI Musings

AGI Musings channel →