Critic Calls Out 'Research Taste Doubles Every 3 Months' Eval as Merely Metric Optimization
dhadfieldmenell · x · 2026-10-06
A widely discussed eval from pzeroresearch claims frontier models' "experimental research taste" has doubled roughly every 3 months since December 2025, with Opus 5.5 exceeding their expert human baseline — a metric they say shapes how fast superintelligence arrives once coding is automated.
Dylan Hadfield-Menell published a pointed green-text critique: after reading the paper, he found the eval actually measures metric optimization (hill-climbing a score) rather than genuine research taste. He says he tried explaining the taste-vs-metric distinction to the authors, who didn't get it and replied "it's a good eval sir." The thread has ignited debate about whether popular AI evals measure what they claim.
Related event: Claim That AI Research Taste Doubles Every 3 Months Draws Skepticism(2 posts)→
More from AGI Musings
- Gary Marcus: OpenAI quietly bolts Python onto LLMs, making the public think they're good at math — Dave_it_up · 2026-10-06
- Why I keep returning to Benedict Evans: 'over time, we change how we work to fit the tool' — _AustinCalvert_ · 2026-10-06
- Expect a wave of 'model welfare' advocacy — driven by persona design, not breakthroughs — alexisgallagher · 2026-10-06
- AI Discourse: 'Capital Realism' and 'Successionism' Are the Same View, and Musk Has Been Circling Them for Years — DavidDuvenaud · 2026-10-06
- Guardian: Altman privatizes AI gains while the public socializes the risks — nordicinst · 2026-10-06
- Huberman: Meta, OpenAI and Anthropic are all becoming biotech companies — Scobleizer · 2026-10-06