Hadfield-Menell: New 'research taste' benchmark measures metric hill-climbing, not taste
dhadfieldmenell · x · 2026-10-07
Responding to Asaf Benj, David Hadfield-Menell concedes the benchmark measures capabilities important for the AI-to-AI R&D feedback loop, but insists words mean something: what it really tests is intuition for hill-climbing a metric, not genuine scientific taste for choosing which problems matter.
Related event: Berkeley Professor Pushes Back on "Research Taste" Benchmark Debate(2 posts)→
More from AGI Musings
- 700+ papers dumped at once: Andy Matuschak calls bulk AI research output 'divine inspiration by the pound' — zetalyrae · 2026-10-07
- Luiza Jarovsky: humanity is about to live in a world of abundant cognition for the first time — LuizaJarovsky · 2026-10-07
- Roman Yampolskiy: the miracle isn't that AI does math, it's that humans ever could — romanyam · 2026-10-07
- Nat Eliason floats an "Alpha art school": academics in the morning, creativity after lunch — nateliason · 2026-10-07
- Neil Chilson to doomers: which factual developments would falsify your beliefs? — neil_chilson · 2026-10-07
- Token efficiency is your security posture: why open models and routing matter for AI defense — sudoraohacker · 2026-10-07