Hadfield-Menell: New 'research taste' benchmark measures metric hill-climbing, not taste

dhadfieldmenell · x · 2026-10-07

Responding to Asaf Benj, David Hadfield-Menell concedes the benchmark measures capabilities important for the AI-to-AI R&D feedback loop, but insists words mean something: what it really tests is intuition for hill-climbing a metric, not genuine scientific taste for choosing which problems matter.

Related event: Berkeley Professor Pushes Back on "Research Taste" Benchmark Debate(2 posts)→

Original post →

More from AGI Musings

AGI Musings channel →