Ofir Press: a good benchmark needs scalable data collection — the hardest step yet

OfirPress · x · 2026-10-02

Princeton researcher Ofir Press outlines three conditions for a good benchmark — an ability you want future AI to have, expressible as verifiable tasks, and challenging for frontier models — plus the hardest fourth: scalable data collection. Benchmarks like SWE-bench, ProgramBench and his new SWE-sweep are great partly because training data can be crawled automatically from the web, which makes benchmark design harder every day.

Original post →

More from Research

Research channel →