Researcher: non-METR AI R&D evals are saturated and uninformative

dfrsrchtwts · x · 2026-09-23

dfrsrchtwts argues (clarifying the claim applies to non-METR evals) that a report section contends AI R&D evals are saturated and uninformative — a view the author shares — while noting METR's Sunlight, Budget NanoGPT, and GamingBot benchmarks may still be somewhat informative.

Related event: Researcher flags false precision in AECI scores and saturated AI R&D benchmarks(5 posts)→

Original post →

More from Models

Models channel →