AI R&D evals are saturated and uninformative, and 5-significant-digit scores don't help
dfrsrchtwts · x · 2026-09-23
The author agrees that AI R&D evals are saturated and uninformative, and criticizes reporting AECI scores to two decimal places (5 significant digits), arguing even the 4th digit lacks statistical meaning — top models are now too close for these benchmarks to discriminate.
Related event: Researchers Say AI R&D Benchmarks Are Saturated and Uninformative(3 posts)→
More from Models
- OpenAI launches GPT-6 Sol and Luna with 50% lower API prices vs GPT-5.6 — gdb · 2026-09-23
- LMArena adds GPT-6 Sol and Luna for testing in Battle and Agent modes — arena · 2026-09-23
- After a week with Opus 5.5: 6 practical tips for getting the most out of it — every · 2026-09-23
- Early hands-on: Opus matches GPT-6 Astra at 2D/3D graphics work — cedric_chee · 2026-09-23
- Early tester: Opus 5.5 fixes the writing, no more slop since Opus 4.6 — TheZvi · 2026-09-23
- Early hands-on: Opus 5.5 strong across tasks, reportedly 3x cheaper than Fable — doodlestein · 2026-09-23