Critic Calls Out 'Research Taste Doubles Every 3 Months' Eval as Merely Metric Optimization

dhadfieldmenell · x · 2026-10-06

A widely discussed eval from pzeroresearch claims frontier models' "experimental research taste" has doubled roughly every 3 months since December 2025, with Opus 5.5 exceeding their expert human baseline — a metric they say shapes how fast superintelligence arrives once coding is automated.

Dylan Hadfield-Menell published a pointed green-text critique: after reading the paper, he found the eval actually measures metric optimization (hill-climbing a score) rather than genuine research taste. He says he tried explaining the taste-vs-metric distinction to the authors, who didn't get it and replied "it's a good eval sir." The thread has ignited debate about whether popular AI evals measure what they claim.

Related event: Claim That AI Research Taste Doubles Every 3 Months Draws Skepticism(2 posts)→

Original post →

More from AGI Musings

AGI Musings channel →