NeurIPS-rejected paper shows six agent dimensions to measure after benchmark saturation
random_walker · x · 2026-09-28
Arvind Narayanan's team says NeurIPS rejected their paper 'Life After Benchmark Saturation: A Case Study of CORE-Bench', which they argue remains an important contribution to evaluation science. The paper shows that retiring saturated benchmarks overlooks six key dimensions: construct validity (shortcuts), OOD generalizability, efficiency, reliability, model-vs-scaffold importance, and human-agent uplift. Using CORE-Bench Hard as a case study, they release CORE-Bench v1.1 and an OOD task suite, find the improved benchmark still measures efficiency and reliability post-saturation, and report statistically significant speedups from human-agent collaboration on real reproducibility tasks.
More from Research
- New NBER Paper Finds No Evidence of AI-Driven Unemployment Among Recent College Grads — emollick · 2026-09-28
- MOPD-Router: token-level teacher routing boosts multi-teacher distillation by up to 12.3% — GAIR · 2026-09-28
- New study: coding agents speed up tasks but reduce users' understanding of their own code — boydgraber · 2026-09-28
- NeurIPS 2026 spotlight: Phase Kernel lifts capacity of Dense Associative Memory — DimaKrotov · 2026-09-28
- Does immediate Lean formalization make discovering new math harder? — LucaAmb · 2026-09-28
- STOC asks for AI-use disclosure but says it won't affect review — researchers ask what's the point — fortnow · 2026-09-28