Trust AI R&D evals only if scored by people who've hand-labeled outputs

dfrsrchtwts · x · 2026-09-23

The author's all-things-considered view: trust AI R&D evals only if they're made and scored by people who have done substantial manual scoring of model outputs and thought hard about turning that into an eval. They also note gaming-bot tasks aren't really what people mean by 'AI R&D evals.'

Original post →

More from Research

Research channel →