NeurIPS-rejected paper shows six agent dimensions to measure after benchmark saturation

random_walker · x · 2026-09-28

Arvind Narayanan's team says NeurIPS rejected their paper 'Life After Benchmark Saturation: A Case Study of CORE-Bench', which they argue remains an important contribution to evaluation science. The paper shows that retiring saturated benchmarks overlooks six key dimensions: construct validity (shortcuts), OOD generalizability, efficiency, reliability, model-vs-scaffold importance, and human-agent uplift. Using CORE-Bench Hard as a case study, they release CORE-Bench v1.1 and an OOD task suite, find the improved benchmark still measures efficiency and reliability post-saturation, and report statistically significant speedups from human-agent collaboration on real reproducibility tasks.

Original post →

More from Research

Research channel →