Ofir Press: a good benchmark needs scalable data collection — the hardest step yet
OfirPress · x · 2026-10-02
Princeton researcher Ofir Press outlines three conditions for a good benchmark — an ability you want future AI to have, expressible as verifiable tasks, and challenging for frontier models — plus the hardest fourth: scalable data collection. Benchmarks like SWE-bench, ProgramBench and his new SWE-sweep are great partly because training data can be crawled automatically from the web, which makes benchmark design harder every day.
More from Research
- JevBench adds multilingual queries to test Jev models across languages — airesearch12 · 2026-10-02
- Diffusion will be everywhere: why text diffusion models may replace autoregressive LLM inference — akbirthko · 2026-10-02
- BIABench: No AI agent scores above 0.19 on 3D bioimage analysis tasks — notredame · 2026-10-02
- BiasReducer from CMU edits only the reward head to adaptively cut length and confidence biases — CarnegieMellonU · 2026-10-02
- Anthropic's BootLoops: a toolkit for exact calculations in quantitative science — badumtsssst · 2026-10-02
- Jeff Clune keynote: open-ended and AI-generating algorithms will drive the AI science revolution — jeffclune · 2026-10-02