Researcher: non-METR AI R&D evals are saturated and uninformative
dfrsrchtwts · x · 2026-09-23
dfrsrchtwts argues (clarifying the claim applies to non-METR evals) that a report section contends AI R&D evals are saturated and uninformative — a view the author shares — while noting METR's Sunlight, Budget NanoGPT, and GamingBot benchmarks may still be somewhat informative.
More from Models
- ProgramBench: rebuilding programs from binaries is brutal — Claude Opus 5 leads at 4.5% resolved — jyangballin · 2026-09-23
- GPT-6 Sol and Luna hit Arena, plus a matched head-to-head vs GPT-5.6 Sol — arena · 2026-09-23
- Plinius leaks full Claude Opus-5.5 system prompt, over 1.9M characters with tools — ivan_bezdomny · 2026-09-23
- Claude Opus 5.5 Frontend Tests: Suspected Quantized fable 5.1, Stable but Heavier Reasoning — karminski3 · 2026-09-23
- Opus 5.5, GPT-6 Sol and Luna drop the same night as AI pace debate rages — ThePeterMick · 2026-09-23
- Travel planner PlanMyVisit switches to GPT-6: faster and half the cost — alexbainbridge · 2026-09-23